Files
sure/test/models/background_job_console_test.rb
Guillem Arias Fauste 5c0a7d8cbd feat(settings): super-admin background jobs console (#2682)
* feat(settings): super-admin background jobs console

Neither managed nor self-hosted production deployments have any view
into background job state (/sidekiq is only mounted outside
production), so a stuck sync, import, or export is invisible until a
user complains — and even then there is no way to act on it.

Adds /settings/background_jobs, gated by the same super-admin
Admin::BaseController as /settings/debug:

- Worker status header: Sidekiq processes, busy count, queue depths
  and latency, retry/scheduled/dead set sizes. Read through a
  fail-closed PORO (BackgroundJobConsole) — when Redis is unreachable
  the console says so and disables actions instead of pretending
  health (deliberate contrast to BackgroundJobHealth's fail-open).
- In-flight operations across all families: incomplete Syncs, Imports
  importing/reverting, ImportSessions importing, FamilyExports
  pending/processing, each with a liveness verdict derived from
  Sidekiq::Workers (job GlobalIDs in running payloads).
- One mutation: mark a presumed-lost operation as such (Sync → stale,
  Import → failed/revert_failed, PdfImport → claim released back to
  pending, ImportSession/FamilyExport → failed).

Guard rails on the mutation, server-side re-checked:

- refused while Redis state is unknown (fail closed)
- refused while the record's job is visibly executing
- refused until the record has been idle past 30 minutes, so a merely
  queued job cannot be shot down and then still run
- refused for parent Syncs with children still in flight (a parent
  legitimately has no live job of its own while children run)
- applied inside with_lock with a status re-check, so a job finishing
  between render and click wins
- audited as a DebugLogEntry (actor, prior status, family)

The cancel endpoint gets its own Rack Attack throttle since the
console deliberately lives under /settings rather than the throttled
/admin prefix (its 10 req/min limit would fight the page's polling).

* fix(jobs-console): count waiting jobs as live and harden cancel paths

Review feedback on #2682 (Codex, CodeRabbit):

- Liveness now covers jobs sitting in queues, the retry set, and the
  scheduled set, not just visibly-executing workers — a job waiting out
  a backlog or retry backoff WILL run later, and most affected job
  classes don't abort on a flipped status, so cancelling invited
  duplicate work. The backlog scan is bounded (5k entries); a truncated
  scan fails closed like redis_error?. The liveness column shows
  "Queued" for these
- find_record! resolves STI subclass names (TransactionImport,
  PdfImport, …) against a base-class whitelist via safe_constantize
  instead of a fixed name map, so non-UI callers naming the subclass
  don't 404
- A PdfImport stuck in reverting goes to revert_failed like every other
  import instead of being released to pending — pending presented a
  possibly half-reverted import as publishable again; only the AI
  extraction claim (importing) is released to pending
- Redis-unreachable warning renders via DS::Alert; operation id cast
  to_s before splitting; admin-cancel error copy moved to i18n

* fix(jobs-console): re-check the stuck window inside the cancel row lock

CodeRabbit round-2: cancellable? evaluates Sidekiq liveness outside the
with_lock transaction (re-running Redis calls under a row lock would be
worse), so a worker picking the job up between the liveness check and
the lock acquisition was invisible. Repeating the updated_at staleness
check inside the lock closes that window — a freshly-started job
touches the record, and the re-check refuses the cancel.

* fix(jobs-console): resolve record_type without reflection, i18n nits

Brakeman flagged safe_constantize on params[:record_type] as a
High-confidence UnsafeReflection (ci/scan_ruby). Replace the
constant lookup with a reverse lookup: find the id in each
cancellable base table and require the claimed type to match the
found record's class or its base class. STI subclass names still
resolve; unknown types still 404.

Also from DS Drift Patrol:
- "Sync · " label in _operation.html.erb now goes through
  t(".sync_label", type:)
- drop the redundant default: on background_jobs_label now that
  the key exists in the locale file
2026-07-21 05:50:40 +02:00

122 lines
4.1 KiB
Ruby

require "test_helper"
class BackgroundJobConsoleTest < ActiveSupport::TestCase
test "fails closed when Sidekiq/Redis is unreachable" do
Sidekiq::ProcessSet.stubs(:new).raises(StandardError.new("no redis"))
console = BackgroundJobConsole.new
sync = Sync.create!(syncable: accounts(:depository), status: :syncing)
sync.update_columns(updated_at: 1.hour.ago)
assert console.redis_error?
assert_nil console.stats
assert_not console.cancellable?(sync)
end
test "detects running records via GlobalIDs in worker payloads" do
sync = Sync.create!(syncable: accounts(:depository), status: :syncing)
sync.update_columns(updated_at: 1.hour.ago)
stub_sidekiq(worker_payloads: [
{ "wrapped" => "SyncJob", "args" => [ { "arguments" => [ { "_aj_globalid" => sync.to_global_id.to_s } ] } ] }
])
console = BackgroundJobConsole.new
assert console.running?(sync)
assert_not console.cancellable?(sync)
end
test "cancellable only when idle past the stuck window with no live job" do
stub_sidekiq
console = BackgroundJobConsole.new
fresh = Sync.create!(syncable: accounts(:depository), status: :syncing)
stuck = Sync.create!(syncable: accounts(:connected), status: :syncing)
stuck.update_columns(updated_at: 1.hour.ago)
assert_not console.cancellable?(fresh)
assert console.cancellable?(stuck)
end
test "a record referenced by a queued job is not cancellable" do
sync = Sync.create!(syncable: accounts(:depository), status: :syncing)
sync.update_columns(updated_at: 1.hour.ago)
stub_sidekiq(queued_items: [
{ "wrapped" => "SyncJob", "args" => [ { "arguments" => [ { "_aj_globalid" => sync.to_global_id.to_s } ] } ] }
])
console = BackgroundJobConsole.new
assert console.enqueued?(sync)
assert_not console.running?(sync)
assert_not console.cancellable?(sync)
end
test "a truncated backlog scan fails closed" do
sync = Sync.create!(syncable: accounts(:depository), status: :syncing)
sync.update_columns(updated_at: 1.hour.ago)
filler = Array.new(BackgroundJobConsole::QUEUE_SCAN_LIMIT + 1) { { "args" => [] } }
stub_sidekiq(queued_items: filler)
console = BackgroundJobConsole.new
assert console.queue_scan_truncated?
assert_not console.cancellable?(sync)
end
test "a parent sync with incomplete children is not cancellable" do
stub_sidekiq
parent = Sync.create!(syncable: families(:dylan_family), status: :syncing)
Sync.create!(syncable: accounts(:depository), status: :syncing, parent: parent)
parent.update_columns(updated_at: 1.hour.ago)
console = BackgroundJobConsole.new
assert_not console.cancellable?(parent)
end
private
def stub_sidekiq(worker_payloads: [], queued_items: [])
process_set = mock("ProcessSet")
process_set.stubs(:size).returns(1)
process_set.stubs(:sum).returns(worker_payloads.size)
Sidekiq::ProcessSet.stubs(:new).returns(process_set)
stats = mock("Stats")
stats.stubs(:enqueued).returns(queued_items.size)
stats.stubs(:retry_size).returns(0)
stats.stubs(:dead_size).returns(0)
stats.stubs(:scheduled_size).returns(0)
Sidekiq::Stats.stubs(:new).returns(stats)
if queued_items.any?
queue = mock("Queue")
queue.stubs(name: "default", size: queued_items.size, latency: 0.0)
queue_yields = queued_items.map { |item| [ stub(item: item) ] }
queue.stubs(:each).multiple_yields(*queue_yields)
Sidekiq::Queue.stubs(:all).returns([ queue ])
else
Sidekiq::Queue.stubs(:all).returns([])
end
empty_set = mock("JobSet")
empty_set.stubs(:each)
Sidekiq::RetrySet.stubs(:new).returns(empty_set)
Sidekiq::ScheduledSet.stubs(:new).returns(empty_set)
workers = mock("Workers")
yields = worker_payloads.map { |payload| [ "process", "thread", { "payload" => payload.to_json } ] }
if yields.any?
workers.stubs(:each).multiple_yields(*yields)
else
workers.stubs(:each)
end
Sidekiq::Workers.stubs(:new).returns(workers)
end
end