Files
sure/app/models/background_job_console.rb
Guillem Arias Fauste 5c0a7d8cbd feat(settings): super-admin background jobs console (#2682)
* feat(settings): super-admin background jobs console

Neither managed nor self-hosted production deployments have any view
into background job state (/sidekiq is only mounted outside
production), so a stuck sync, import, or export is invisible until a
user complains — and even then there is no way to act on it.

Adds /settings/background_jobs, gated by the same super-admin
Admin::BaseController as /settings/debug:

- Worker status header: Sidekiq processes, busy count, queue depths
  and latency, retry/scheduled/dead set sizes. Read through a
  fail-closed PORO (BackgroundJobConsole) — when Redis is unreachable
  the console says so and disables actions instead of pretending
  health (deliberate contrast to BackgroundJobHealth's fail-open).
- In-flight operations across all families: incomplete Syncs, Imports
  importing/reverting, ImportSessions importing, FamilyExports
  pending/processing, each with a liveness verdict derived from
  Sidekiq::Workers (job GlobalIDs in running payloads).
- One mutation: mark a presumed-lost operation as such (Sync → stale,
  Import → failed/revert_failed, PdfImport → claim released back to
  pending, ImportSession/FamilyExport → failed).

Guard rails on the mutation, server-side re-checked:

- refused while Redis state is unknown (fail closed)
- refused while the record's job is visibly executing
- refused until the record has been idle past 30 minutes, so a merely
  queued job cannot be shot down and then still run
- refused for parent Syncs with children still in flight (a parent
  legitimately has no live job of its own while children run)
- applied inside with_lock with a status re-check, so a job finishing
  between render and click wins
- audited as a DebugLogEntry (actor, prior status, family)

The cancel endpoint gets its own Rack Attack throttle since the
console deliberately lives under /settings rather than the throttled
/admin prefix (its 10 req/min limit would fight the page's polling).

* fix(jobs-console): count waiting jobs as live and harden cancel paths

Review feedback on #2682 (Codex, CodeRabbit):

- Liveness now covers jobs sitting in queues, the retry set, and the
  scheduled set, not just visibly-executing workers — a job waiting out
  a backlog or retry backoff WILL run later, and most affected job
  classes don't abort on a flipped status, so cancelling invited
  duplicate work. The backlog scan is bounded (5k entries); a truncated
  scan fails closed like redis_error?. The liveness column shows
  "Queued" for these
- find_record! resolves STI subclass names (TransactionImport,
  PdfImport, …) against a base-class whitelist via safe_constantize
  instead of a fixed name map, so non-UI callers naming the subclass
  don't 404
- A PdfImport stuck in reverting goes to revert_failed like every other
  import instead of being released to pending — pending presented a
  possibly half-reverted import as publishable again; only the AI
  extraction claim (importing) is released to pending
- Redis-unreachable warning renders via DS::Alert; operation id cast
  to_s before splitting; admin-cancel error copy moved to i18n

* fix(jobs-console): re-check the stuck window inside the cancel row lock

CodeRabbit round-2: cancellable? evaluates Sidekiq liveness outside the
with_lock transaction (re-running Redis calls under a row lock would be
worse), so a worker picking the job up between the liveness check and
the lock acquisition was invisible. Repeating the updated_at staleness
check inside the lock closes that window — a freshly-started job
touches the record, and the re-check refuses the cancel.

* fix(jobs-console): resolve record_type without reflection, i18n nits

Brakeman flagged safe_constantize on params[:record_type] as a
High-confidence UnsafeReflection (ci/scan_ruby). Replace the
constant lookup with a reverse lookup: find the id in each
cancellable base table and require the claimed type to match the
found record's class or its base class. STI subclass names still
resolve; unknown types still 404.

Also from DS Drift Patrol:
- "Sync · " label in _operation.html.erb now goes through
  t(".sync_label", type:)
- drop the redundant default: on background_jobs_label now that
  the key exists in the locale file
2026-07-21 05:50:40 +02:00

181 lines
6.3 KiB
Ruby

require "sidekiq/api"
# Backs the super-admin background jobs console (/settings/background_jobs).
#
# Combines domain truth (in-flight Sync / Import / ImportSession / FamilyExport
# records across all families) with Sidekiq runtime truth (worker processes,
# queue depths, jobs currently executing) so an operator can tell a running
# job from one whose worker died and mark the latter as lost.
#
# Unlike BackgroundJobHealth this fails CLOSED: when Redis is unreachable,
# liveness is unknown and destructive actions are refused rather than the
# console pretending everything is healthy.
class BackgroundJobConsole
OPERATIONS_LIMIT = 100
# A record younger than this may belong to a job that is merely queued
# behind a backlog or between Sidekiq heartbeats — refuse to touch it.
STUCK_AFTER = 30.minutes
# Upper bound on queue/retry/schedule entries scanned for record
# references. Past this the backlog is inspected only partially, so
# liveness is unknowable and cancellation fails closed (like redis_error?).
QUEUE_SCAN_LIMIT = 5_000
Stats = Struct.new(:processes, :busy, :enqueued, :retry_size, :dead_size, :scheduled_size, :queues, keyword_init: true)
attr_reader :stats
def initialize
@redis_error = false
@queue_scan_truncated = false
@running_global_ids = Set.new
@queued_global_ids = Set.new
@stats = nil
load_runtime_state
end
def redis_error?
@redis_error
end
def queue_scan_truncated?
@queue_scan_truncated
end
# In-flight operations across ALL families, newest activity first. This is
# deliberately instance-global — the console is super-admin only.
def operations
@operations ||= [
Sync.incomplete.includes(:syncable).order(updated_at: :desc).limit(OPERATIONS_LIMIT).to_a,
Import.where(status: [ :importing, :reverting ]).includes(:family).order(updated_at: :desc).limit(OPERATIONS_LIMIT).to_a,
ImportSession.where(status: :importing).includes(:family).order(updated_at: :desc).limit(OPERATIONS_LIMIT).to_a,
FamilyExport.where(status: [ :pending, :processing ]).includes(:family).order(updated_at: :desc).limit(OPERATIONS_LIMIT).to_a
].flatten.sort_by(&:updated_at).reverse
end
# A record's job is visibly executing right now (its GlobalID appears in a
# worker's payload). False also means "unknown" when redis_error? is set —
# callers must check that first for destructive decisions.
def running?(record)
@running_global_ids.include?(record.to_global_id.to_s)
end
# A job referencing this record is sitting in a queue, the retry set, or
# the scheduled set. Such a job WILL run later — most of the affected job
# classes don't abort just because an operator flipped the record's status,
# so terminalizing now would invite duplicate/conflicting work when it
# finally executes.
def enqueued?(record)
@queued_global_ids.include?(record.to_global_id.to_s)
end
# Safe to force-terminalize: liveness is knowable (Redis reachable, backlog
# scan complete), no job referencing the record is executing or waiting to
# execute, the record has been idle past the stuck window, and (for syncs)
# no children are still in flight — a parent Sync legitimately has no live
# job of its own while its children run.
def cancellable?(record)
return false if redis_error?
return false if queue_scan_truncated?
return false if running?(record)
return false if enqueued?(record)
return false if record.updated_at > STUCK_AFTER.ago
return false if record.is_a?(Sync) && record.children.incomplete.exists?
true
end
def self.family_for(record)
if record.is_a?(Sync)
syncable = record.syncable
syncable.is_a?(Family) ? syncable : syncable&.family
else
record.family
end
end
private
def load_runtime_state
processes = Sidekiq::ProcessSet.new
sidekiq_stats = Sidekiq::Stats.new
@stats = Stats.new(
processes: processes.size,
busy: processes.sum { |process| process["busy"].to_i },
enqueued: sidekiq_stats.enqueued,
retry_size: sidekiq_stats.retry_size,
dead_size: sidekiq_stats.dead_size,
scheduled_size: sidekiq_stats.scheduled_size,
queues: Sidekiq::Queue.all.map { |queue| { name: queue.name, size: queue.size, latency: queue.latency.round(1) } }
)
@running_global_ids = collect_running_global_ids
@queued_global_ids = collect_queued_global_ids
rescue => e
Rails.logger.warn("BackgroundJobConsole: Sidekiq state unavailable: #{e.class}: #{e.message}")
@redis_error = true
@stats = nil
@running_global_ids = Set.new
@queued_global_ids = Set.new
end
# All jobs are ActiveJob-wrapped, so record references appear in worker
# payloads as serialized GlobalIDs ({"_aj_globalid" => "gid://..."}).
def collect_running_global_ids
ids = Set.new
Sidekiq::Workers.new.each do |_process_id, _thread_id, work|
payload = work.respond_to?(:payload) ? work.payload : work["payload"]
payload = JSON.parse(payload) if payload.is_a?(String)
collect_global_ids(payload, ids)
rescue JSON::ParserError
next
end
ids
end
# Record references in jobs that are waiting to run: queue backlogs, the
# retry set, and the scheduled set. Bounded by QUEUE_SCAN_LIMIT — on a
# truncated scan, cancellation fails closed via queue_scan_truncated?.
def collect_queued_global_ids
ids = Set.new
scanned = 0
each_waiting_job do |item|
scanned += 1
if scanned > QUEUE_SCAN_LIMIT
@queue_scan_truncated = true
break
end
collect_global_ids(item, ids)
end
ids
end
def each_waiting_job(&block)
Sidekiq::Queue.all.each do |queue|
queue.each { |job| yield job.item }
end
Sidekiq::RetrySet.new.each { |job| yield job.item }
Sidekiq::ScheduledSet.new.each { |job| yield job.item }
end
def collect_global_ids(node, ids)
case node
when Hash
node.each do |key, value|
if key == "_aj_globalid" && value.is_a?(String)
ids << value
else
collect_global_ids(value, ids)
end
end
when Array
node.each { |value| collect_global_ids(value, ids) }
end
end
end