Repository navigation
Outbox relay strands events in processing state after a crash #417
Description
Activity
- addedbugSomething isn't workingSomething isn't workingdatabaseImported from PRODUCTION_ISSUES.mdImported from PRODUCTION_ISSUES.mdreliabilityImported from PRODUCTION_ISSUES.mdImported from PRODUCTION_ISSUES.md
on Aug 20, 2026 Hi team,
I'm applying to work on this issue.
I have strong experience with TypeScript, NestJS, and Soroban smart contracts. I will deliver clean, tested code following the acceptance criteria.
Ready to start immediately.
Best regards,
samkay-opsHi team,
I'm applying to work on this issue.
I have strong experience with TypeScript, NestJS, and Soroban smart contracts. I will deliver clean, tested code following the acceptance criteria.
Ready to start immediately.
Best regards,
samkay-ops- addedGrantFox OSSIssue tracked in GrantFox OSSIssue tracked in GrantFox OSSMaybe RewardedIssue may be eligible for a GrantFox rewardIssue may be eligible for a GrantFox rewardThird CampaignCampaign: Third CampaignCampaign: Third Campaign
on Aug 20, 2026 temiport25 commented
on Aug 20, 2026 ContributorMore actionscan i solve this?
grantfox-oss commented
on Aug 20, 2026 grantfox-ossboton Aug 20, 2026 – with GrantFox OSSMore actions🦊 GrantFox — @temiport25 has been assigned to this issue as part of the Third Campaign campaign!
Next steps:
- Open a Pull Request referencing this issue (e.g.,
Closes #417) - Your PR will be reviewed by the Quantarq maintainers
Good luck! Track your progress on GrantFox.
- Open a Pull Request referencing this issue (e.g.,
grantfox-oss commented
on Aug 22, 2026 grantfox-ossboton Aug 22, 2026 – with GrantFox OSSMore actions🎉 This issue has been marked as completed on GrantFox as part of the Third Campaign campaign!
@temiport25's PR #432 was approved and merged by @YaronZaki.
🏆 @temiport25: You earned 35 FoxPoints for this contribution! Your current tier: Explorer (745 total points). Track your full progress on GrantFox.
👏 Great work, @temiport25! Keep contributing to Quantarq.
Labels / Complexity: bug, reliability, database, Backend · Extremely High — 500
Problem
OutboxRelay.process_pending_eventsinquantara/web_app/tasks/outbox_relay.pymarks an event"processing"and commits it before publishing to Celery:The re-scan query only selects
status.in_(["pending", "failed"]), so an event stuck in"processing"is never picked up again. A crash, broker outage, or exception betweendb.commit()anddelay()(or adelay()that raises because Redis/Celery is down) permanently strands the event: it will never be re-queued, never fail, and never be retried, even thoughretry_countremains belowmax_retries.Secondary defects in the same path:
except Exceptioncatches the faileddelay()but only logs; it does not revert the event to"pending"or"failed", so the stranding window is the default outcome on any publish error.process_position_opened_taskcallsposition_db_connector.get_object(Position, position_id)whereposition_idis the raw JSON string from the payload;Position.idis a PostgreSQLUUIDcolumn (quantara/web_app/db/models.py), so the comparison can raise aDataErroron invalid UUID syntax and skip legitimately processed events.Root cause
Why this is architecturally hard
"pending"in theexcept— only narrows the window; it does not eliminate the crash-between-commit-and-publish gap. A correct fix needs an atomic claim (e.g. a conditionalUPDATE ... WHERE status='pending' RETURNING idwith a claim timestamp, or moving the enqueue before the commit and rolling back on failure).SessionLocal()sessions and run in different processes, so there is no shared transaction to make the mark-and-enqueue atomic; the contributor must introduce an explicit lease/claim mechanism rather than reusedb.commit().processingwithupdated_at/claimed_at, and re-claim events whose lease has expired), which changes the re-scan query and needs a backoff so hot events are not double-queued.get_objectmust be fixed or the task's payload validated, otherwise the reliability fix will still fail at the resource level.Proposed design
Atomic claim with lease expiry:
and a re-scan that reclaims
processingevents older than a lease timeout. Offer this as one option; the maintainer decision is the lease duration and whether to use Celery's own retry or the outboxretry_count.Downstream impact
No public API change. The
OutboxEventtable (quantara/web_app/db/models.py) may need aclaimed_atcolumn, which is an Alembic migration. Existing rows stuck inprocessingmust be handled by the migration or a one-off reclaim.Acceptance criteria
Service
delay()between claim and enqueue does not permanently strand an event; the event is reclaimed and re-queued within a bounded time.processingevents.process_position_opened_taskparses/validates theposition_idbefore the UUID comparison.Tests
processingevent and assert it is reclaimed after the lease expires.cd quantara && poetry run pytest web_app/tests.Out of scope
Do not rebuild the Celery wiring or add a dead-letter store in this issue; only make the outbox claim/reclaim durable.
Getting started
Files in scope:
quantara/web_app/tasks/outbox_relay.py,quantara/web_app/db/models.py, plus an Alembic migration. Verify with:Good first files to read:
quantara/web_app/tasks/outbox_relay.py,quantara/web_app/db/models.py(theOutboxEventmodel),quantara/web_app/db/crud/position.py.