Files
sure/app/models/sync.rb
Guillem Arias Fauste 5105752bbd feat(sync): family-facing cancellation for syncs, imports, and exports (#2685)
* feat(sync): family-facing cancellation for syncs, imports, and exports

Users had no way to stop or recover any background operation: a
mistaken "Sync all" runs to completion, and an import or export whose
job died (hard worker kills lose in-flight Sidekiq jobs) wedges with a
spinner forever.

Sync cancellation (cooperative — nothing is ever killed):

- New syncs.cancel_requested_at column. Only the cancelled sync
  carries the flag: pending descendants are marked stale immediately
  (their queued jobs no-op via the existing may_start? guard), while
  descendants whose jobs are already executing finish their work
  honestly.
- Family::Syncer stops fanning out child syncs once the flag is set
  (fresh read per iteration — the flag comes from the web process).
- Finalization resolves a cancel-requested sync to stale instead of
  completed, which also skips post-sync (transfer matching, rules,
  broadcasts) via the existing stale gate.
- The `visible` scope excludes cancel-requested syncs, so spinners
  clear immediately and — fixing a latent bug this feature would have
  amplified — sync_later no longer piggybacks a new sync request onto
  a dying sync it would silently swallow.
- Cancel button appears next to "Sync all" on the accounts page while
  a family sync is visible. SyncsController#cancel scopes through
  Sync.for_family with resource_owner, so cross-family ids 404 and
  account-level syncs respect per-user account access.

Stuck import/export self-service:

- Import#force_fail! / FamilyExport#force_fail!: allowed only once the
  record has been idle past PRESUMED_LOST_AFTER (1 hour — dwarfs any
  legitimate run), and applied inside with_lock with a status
  re-check, so a job finishing between page render and button click
  wins. Imports fail into the existing retry path (reverting ->
  revert_failed keeps the revert retryable); PdfImports release their
  processing claim back to pending; exports fail so a new one can be
  created.
- "Mark as failed" buttons appear on the imports/exports index rows
  only when a record is presumed lost, behind the pages' existing
  permission gates (statement-import permission for imports, admin
  for exports).

* fix(sync-cancel): cascade pending cancels, guard late finalizers, scope provider syncs

Review feedback on #2685 (CodeRabbit, Codex):

- request_cancel! now cascades finalization for pending syncs too: a
  pending child resolved to stale never runs its job, so nothing else
  would ever call finalize_if_all_children_finalized — its waiting
  parent hung in syncing until the 24h sweep (CodeRabbit critical)
- SimplefinItem::Syncer#mark_completed re-reads the sync under a row
  lock and skips finalization once cancellation was requested or the
  row went terminal — its in-memory copy predates the cancel, and the
  unguarded complete! (plus the raw status fallback) resurrected a
  cancelled sync and re-ran post-sync (Codex)
- Cancelling provider-item syncs now requires admin: for_family's
  resource_owner only scopes the Account branch, so a restricted member
  could cancel admin-managed provider syncs spanning accounts they
  cannot see. Family- and account-level syncs stay member-cancellable,
  matching the buttons the UI shows (Codex)
- Lost-import error copy moved behind i18n
  (imports.errors.presumed_lost), resolved at call time (CodeRabbit)
- Tests: pending-child cancel finalizes the parent; late provider
  complete! cannot resurrect a cancelled sync; provider-sync
  cancellation is admin-only

* fix(sync-cancel): capture the skipped-finalization case via DebugLogEntry

CodeRabbit round-2: the mark_completed skip (cancelled/terminal sync)
is support-relevant — record it in the super-admin debug UI with the
family and provider attached instead of a raw Rails.logger line.
category: provider_sync, matching the other provider syncers.
2026-07-17 06:57:54 +02:00

342 lines
11 KiB
Ruby

class Sync < ApplicationRecord
# We run a cron that marks any syncs that have not been resolved in 24 hours as "stale"
# Syncs often become stale when new code is deployed and the worker restarts
STALE_AFTER = 24.hours
# The max time that a sync will show in the UI (after 5 minutes)
VISIBLE_FOR = 5.minutes
include AASM
Error = Class.new(StandardError)
belongs_to :syncable, polymorphic: true
belongs_to :parent, class_name: "Sync", optional: true
has_many :children, class_name: "Sync", foreign_key: :parent_id, dependent: :destroy
scope :ordered, -> { order(created_at: :desc, id: :desc) }
scope :incomplete, -> { where("syncs.status IN (?)", %w[pending syncing]) }
# Cancel-requested syncs are excluded so spinners clear immediately and
# sync_later stops piggybacking new requests onto a dying sync.
scope :visible, -> { incomplete.where("syncs.created_at > ?", VISIBLE_FOR.ago).where(cancel_requested_at: nil) }
after_commit :update_family_sync_timestamp, on: [ :create, :update ]
serialize :sync_stats, coder: JSON
validate :window_valid
# Sync state machine
aasm column: :status, timestamps: true do
state :pending, initial: true
state :syncing
state :completed
state :failed
state :stale
after_all_transitions :handle_transition
event :start, after_commit: :handle_start_transition do
transitions from: :pending, to: :syncing
end
event :complete, after_commit: :handle_completion_transition do
transitions from: :syncing, to: :completed
end
event :fail do
transitions from: :syncing, to: :failed
end
# Marks a sync that never completed within the expected time window
event :mark_stale do
transitions from: %i[pending syncing], to: :stale
end
end
class << self
def clean
incomplete.where("syncs.created_at < ?", STALE_AFTER.ago).find_each(&:mark_stale!)
end
def for_family(family, resource_owner: nil)
query = where(syncable_type: "Family", syncable_id: family.id)
query = query.or(where(syncable_type: "Account", syncable_id: account_syncable_ids(family, resource_owner)))
family_syncable_associations.each do |association|
query = query.or(
where(syncable_type: association.klass.name, syncable_id: family.public_send(association.name).select(:id))
)
end
query
end
# True iff the family has any pending/syncing Sync — across its own row,
# its accounts, and every Syncable provider `*_items` association. Built
# on `for_family` so new provider integrations are picked up automatically
# via `family_syncable_associations` reflection (no hand-rolled list).
def any_incomplete_for?(family)
for_family(family).incomplete.exists?
end
private
def account_syncable_ids(family, resource_owner)
(resource_owner ? resource_owner.accessible_accounts : family.accounts)
.where(family_id: family.id)
.select(:id)
end
def family_syncable_associations
Family.reflect_on_all_associations(:has_many).select do |association|
association.name.to_s.end_with?("_items") &&
association.klass.included_modules.include?(Syncable)
rescue NameError
false
end
end
end
def in_progress?
pending? || syncing?
end
def terminal?
completed? || failed? || stale?
end
def api_error_payload
return unless failed? || stale?
return if stale? && error.blank?
{
message: stale? ? "Sync became stale before completion" : "Sync failed"
}
end
def perform
Rails.logger.tagged("Sync", id, syncable_type, syncable_id) do
# This can happen on server restarts or if Sidekiq enqueues a duplicate job
unless may_start?
Rails.logger.warn("Sync #{id} is not in a valid state (#{aasm.from_state}) to start. Skipping sync.")
return
end
# Guard: syncable may have been deleted while job was queued
unless syncable.present?
Rails.logger.warn("Sync #{id} - syncable #{syncable_type}##{syncable_id} no longer exists. Marking as failed.")
start! if may_start?
fail! if may_fail?
update(error: "Syncable record was deleted")
return
end
# Guard: syncable may be scheduled for deletion
if syncable.respond_to?(:scheduled_for_deletion?) && syncable.scheduled_for_deletion?
Rails.logger.warn("Sync #{id} - syncable #{syncable_type}##{syncable_id} is scheduled for deletion. Skipping sync.")
start! if may_start?
fail! if may_fail?
update(error: "Syncable record is scheduled for deletion")
return
end
start!
begin
syncable.perform_sync(self)
rescue => e
# Re-check state under a row lock (with_lock reloads): the sync may
# have been terminalized externally (marked stale by SyncCleanerJob)
# while this job was still running. An unguarded fail! on the in-memory
# record would silently overwrite that terminal status.
with_lock { fail! if may_fail? }
update(error: e.message)
report_error(e)
ensure
finalize_if_all_children_finalized
end
end
end
# Requests cooperative cancellation of this sync tree. Only this sync
# carries the flag: pending descendants are marked stale immediately (their
# queued jobs no-op via the may_start? guard), while descendants whose jobs
# are already running finish their work honestly — finalization then
# resolves this sync to stale instead of completed. Returns false when the
# sync is already terminal.
def request_cancel!
result = with_lock do
if pending?
# Job hasn't started — safe to resolve immediately; the queued job
# will no-op via the may_start? guard.
update!(cancel_requested_at: Time.current)
mark_stale!
:cancelled_before_start
elsif syncing?
update!(cancel_requested_at: Time.current)
:cancel_requested
end
end
return false if result.nil?
# Both paths cascade: a pending sync resolved above went terminal without
# its job ever running, so nothing else will ever call
# finalize_if_all_children_finalized for it — without this, a parent
# waiting on the cancelled child stays syncing until the 24h sweep.
# finalize_if_all_children_finalized re-reads under lock!, so it safely
# no-ops on this sync's own branch when already terminal.
cancel_pending_descendants!
finalize_if_all_children_finalized
true
end
# Fresh DB read — cancellation is requested from the web process while this
# sync's job holds a stale in-memory copy of the record.
def cancel_requested?
self.class.where(id: id).pick(:cancel_requested_at).present?
end
# Finalizes the current sync AND parent (if it exists)
def finalize_if_all_children_finalized
Sync.transaction do
lock!
# Eagerly load children once so that all_children_finalized? and
# has_failed_children? can filter in-memory without additional DB queries.
children.load
# If this is the "parent" and there are still children running, don't finalize.
return unless all_children_finalized?
if syncing?
if cancel_requested_at?
# User asked for cancellation while work was in flight. Whatever
# children completed keep their data; the tree resolves to stale
# (which also skips post-sync below).
mark_stale!
elsif has_failed_children?
fail!
else
complete!
end
end
# If we make it here, the sync is finalized. Run post-sync, regardless of failure/success —
# unless the sync was terminalized externally (marked stale by SyncCleanerJob while its job
# was still running). A stale sync's job has been written off: re-running transfer matching,
# rules, and broadcasts for it would apply side effects for work the system already abandoned.
perform_post_sync unless stale?
end
# If this sync has a parent, try to finalize it so the child status propagates up the chain.
parent&.finalize_if_all_children_finalized
end
# If a sync is pending, we can adjust the window if new syncs are created with a wider window.
def expand_window_if_needed(new_window_start_date, new_window_end_date)
return unless pending?
return if self.window_start_date.nil? && self.window_end_date.nil? # already as wide as possible
earliest_start_date = if self.window_start_date && new_window_start_date
[ self.window_start_date, new_window_start_date ].min
else
nil
end
latest_end_date = if self.window_end_date && new_window_end_date
[ self.window_end_date, new_window_end_date ].max
else
nil
end
update(
window_start_date: earliest_start_date,
window_end_date: latest_end_date
)
end
protected
def cancel_pending_descendants!
children.incomplete.find_each do |child|
child.with_lock { child.mark_stale! if child.pending? }
child.cancel_pending_descendants!
end
end
private
def log_status_change
Rails.logger.info("changing from #{aasm.from_state} to #{aasm.to_state} (event: #{aasm.current_event})")
end
def has_failed_children?
children.any?(&:failed?)
end
def all_children_finalized?
children.none? { |child| child.pending? || child.syncing? }
end
def perform_post_sync
Rails.logger.info("Performing post-sync for #{syncable_type} (#{syncable.id})")
syncable.perform_post_sync
syncable.broadcast_sync_complete
rescue => e
Rails.logger.error("Error performing post-sync for #{syncable_type} (#{syncable.id}): #{e.message}")
report_error(e)
end
def report_error(error)
Sentry.capture_exception(error) do |scope|
scope.set_tags(sync_id: id)
end
end
def report_warnings
todays_sync_count = syncable.syncs.where(created_at: Date.current.all_day).count
if todays_sync_count > 10
Sentry.capture_exception(
Error.new("#{syncable_type} (#{syncable.id}) has exceeded 10 syncs today (count: #{todays_sync_count})"),
level: :warning
)
end
end
def handle_start_transition
report_warnings
end
def handle_transition
log_status_change
end
def handle_completion_transition
family.touch(:latest_sync_completed_at)
end
def window_valid
if window_start_date && window_end_date && window_start_date > window_end_date
errors.add(:window_end_date, "must be greater than window_start_date")
end
end
def update_family_sync_timestamp
return if syncable.nil?
return unless family&.persisted?
family.touch(:latest_sync_activity_at)
end
def family
return nil unless syncable
if syncable.is_a?(Family)
syncable
else
syncable.family
end
end
end