mirror of
https://github.com/we-promise/sure.git
synced 2026-08-12 11:10:20 +00:00
* fix(chat): surface and recover from undelivered assistant responses When the background worker that runs AssistantResponseJob is down or not polling the high_priority queue, the eager "pending" assistant message had no safeguard: chat hung on "Thinking…" forever with no error, timeout, or retry, and the LLM provider was never even called. Add three layers of resilience: - Client watchdog (chat_controller.js): a pending "Thinking…" bubble that waits past a threshold (default 90s) with no response asks the server to fail it. Keyed on the pending marker — which disappears the instant a real response streams — so a slow-but-working response is never falsely timed out. - Server failure capture (Chat#handle_undelivered_response! + MessagesController#report_timeout): clears the dead bubble, records a friendly "the assistant didn't respond" error with Retry, and writes a DebugLogEntry so support can see it in /settings/debug. - Worker liveness signal (BackgroundJobHealth + warning banner): reads Sidekiq's process/queue state directly from Redis in the web process — so a down worker is detectable even though the worker is the thing that's broken — and warns self-hosted users when no worker polls high_priority or it's badly backed up. Fails open so a Sidekiq/Redis blip never blocks chat. Tests: Chat model (3), MessagesController (2), BackgroundJobHealth (5). * fix(chat): address review — server-side timeout guard + watchdog retry safety - Chat#handle_undelivered_response!: gate the state change behind a row lock and a server-side minimum age (UNDELIVERED_RESPONSE_TIMEOUT = 60s). The browser watchdog is untrusted, so re-read the row under lock and only fail a bubble that is still pending AND has genuinely waited past the timeout — never racing a worker that is finishing a slow response. (Codex P2 + CodeRabbit) - chat_controller.js: only mark a report URL as reported on response.ok (fetch resolves on HTTP 4xx/5xx, rejecting only on network errors), and guard concurrent duplicate POSTs with an in-flight set, so a failed report retries instead of stranding the bubble. (CodeRabbit) - _worker_health_warning: drop the unused `chat:` local + its render arg. (CodeRabbit)
51 lines
1.7 KiB
Ruby
51 lines
1.7 KiB
Ruby
require "test_helper"
|
|
|
|
class BackgroundJobHealthTest < ActiveSupport::TestCase
|
|
setup { Rails.cache.delete(BackgroundJobHealth::CACHE_KEY) }
|
|
teardown { Rails.cache.delete(BackgroundJobHealth::CACHE_KEY) }
|
|
|
|
test "healthy when a worker polls the critical queue with low latency" do
|
|
stub_sidekiq(processes: [ { "queues" => [ "high_priority", "default" ] } ], latency: 1.0)
|
|
|
|
assert BackgroundJobHealth.healthy?
|
|
assert_equal 1, BackgroundJobHealth.snapshot[:workers]
|
|
end
|
|
|
|
test "unhealthy when no workers are online" do
|
|
stub_sidekiq(processes: [], latency: 0.0)
|
|
|
|
assert_not BackgroundJobHealth.healthy?
|
|
end
|
|
|
|
test "unhealthy when the critical queue is not polled by any worker" do
|
|
stub_sidekiq(processes: [ { "queues" => [ "default" ] } ], latency: 0.0)
|
|
|
|
assert_not BackgroundJobHealth.healthy?
|
|
end
|
|
|
|
test "unhealthy when the critical queue is badly backed up" do
|
|
stub_sidekiq(processes: [ { "queues" => [ "high_priority" ] } ], latency: 9_999)
|
|
|
|
assert_not BackgroundJobHealth.healthy?
|
|
end
|
|
|
|
test "fails open when Sidekiq/Redis is unreachable" do
|
|
Sidekiq::ProcessSet.stubs(:new).raises(StandardError.new("no redis"))
|
|
|
|
assert BackgroundJobHealth.healthy?
|
|
assert BackgroundJobHealth.snapshot[:error]
|
|
end
|
|
|
|
private
|
|
def stub_sidekiq(processes:, latency:)
|
|
process_set = mock("ProcessSet")
|
|
process_set.stubs(:size).returns(processes.size)
|
|
process_set.stubs(:flat_map).returns(processes.flat_map { |p| p["queues"] })
|
|
Sidekiq::ProcessSet.stubs(:new).returns(process_set)
|
|
|
|
queue = mock("Queue")
|
|
queue.stubs(:latency).returns(latency)
|
|
Sidekiq::Queue.stubs(:new).with("high_priority").returns(queue)
|
|
end
|
|
end
|