Files
sure/test/models/background_job_health_test.rb
T
Guillem Arias FausteandGitHub 169fd139b5 fix(chat): surface and recover from undelivered assistant responses (#2436)
* fix(chat): surface and recover from undelivered assistant responses

When the background worker that runs AssistantResponseJob is down or not
polling the high_priority queue, the eager "pending" assistant message had no
safeguard: chat hung on "Thinking…" forever with no error, timeout, or retry,
and the LLM provider was never even called.

Add three layers of resilience:

- Client watchdog (chat_controller.js): a pending "Thinking…" bubble that waits
  past a threshold (default 90s) with no response asks the server to fail it.
  Keyed on the pending marker — which disappears the instant a real response
  streams — so a slow-but-working response is never falsely timed out.

- Server failure capture (Chat#handle_undelivered_response! +
  MessagesController#report_timeout): clears the dead bubble, records a friendly
  "the assistant didn't respond" error with Retry, and writes a DebugLogEntry so
  support can see it in /settings/debug.

- Worker liveness signal (BackgroundJobHealth + warning banner): reads Sidekiq's
  process/queue state directly from Redis in the web process — so a down worker
  is detectable even though the worker is the thing that's broken — and warns
  self-hosted users when no worker polls high_priority or it's badly backed up.
  Fails open so a Sidekiq/Redis blip never blocks chat.

Tests: Chat model (3), MessagesController (2), BackgroundJobHealth (5).

* fix(chat): address review — server-side timeout guard + watchdog retry safety

- Chat#handle_undelivered_response!: gate the state change behind a row lock and
  a server-side minimum age (UNDELIVERED_RESPONSE_TIMEOUT = 60s). The browser
  watchdog is untrusted, so re-read the row under lock and only fail a bubble
  that is still pending AND has genuinely waited past the timeout — never racing
  a worker that is finishing a slow response. (Codex P2 + CodeRabbit)
- chat_controller.js: only mark a report URL as reported on response.ok (fetch
  resolves on HTTP 4xx/5xx, rejecting only on network errors), and guard
  concurrent duplicate POSTs with an in-flight set, so a failed report retries
  instead of stranding the bubble. (CodeRabbit)
- _worker_health_warning: drop the unused `chat:` local + its render arg. (CodeRabbit)
2026-06-25 05:44:44 +02:00

51 lines
1.7 KiB
Ruby

require "test_helper"
class BackgroundJobHealthTest < ActiveSupport::TestCase
setup { Rails.cache.delete(BackgroundJobHealth::CACHE_KEY) }
teardown { Rails.cache.delete(BackgroundJobHealth::CACHE_KEY) }
test "healthy when a worker polls the critical queue with low latency" do
stub_sidekiq(processes: [ { "queues" => [ "high_priority", "default" ] } ], latency: 1.0)
assert BackgroundJobHealth.healthy?
assert_equal 1, BackgroundJobHealth.snapshot[:workers]
end
test "unhealthy when no workers are online" do
stub_sidekiq(processes: [], latency: 0.0)
assert_not BackgroundJobHealth.healthy?
end
test "unhealthy when the critical queue is not polled by any worker" do
stub_sidekiq(processes: [ { "queues" => [ "default" ] } ], latency: 0.0)
assert_not BackgroundJobHealth.healthy?
end
test "unhealthy when the critical queue is badly backed up" do
stub_sidekiq(processes: [ { "queues" => [ "high_priority" ] } ], latency: 9_999)
assert_not BackgroundJobHealth.healthy?
end
test "fails open when Sidekiq/Redis is unreachable" do
Sidekiq::ProcessSet.stubs(:new).raises(StandardError.new("no redis"))
assert BackgroundJobHealth.healthy?
assert BackgroundJobHealth.snapshot[:error]
end
private
def stub_sidekiq(processes:, latency:)
process_set = mock("ProcessSet")
process_set.stubs(:size).returns(processes.size)
process_set.stubs(:flat_map).returns(processes.flat_map { |p| p["queues"] })
Sidekiq::ProcessSet.stubs(:new).returns(process_set)
queue = mock("Queue")
queue.stubs(:latency).returns(latency)
Sidekiq::Queue.stubs(:new).with("high_priority").returns(queue)
end
end