From f78303ebbebfb6fb96a7d654701bfcb0bd048ed6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Juan=20Jos=C3=A9=20Mata?= Date: Mon, 31 Aug 2026 10:47:48 -0700 Subject: [PATCH] Add function-calling probe to detect tool-use support (#3255) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit * feat(ai-health): name the missing function calling behind an opaque chat error The assistant reads accounts, transactions and holdings through function calls, so every chat request carries a `tools` payload. A model without function-calling support rejects it — OpenRouter answers a bare 404 — and the operator sees only that status code, with nothing pointing at the model. Both earlier attempts at this guessed from the chat-time error; the AI status page already runs live probes, so let it answer the question directly instead. `AiHealth::Probe#function_calling` asks the configured model for one trivial tool call the way the assistant asks for its own: chat completions with `tools` for OpenAI-compatible endpoints, the Responses API for hosted OpenAI, and `messages.create` with `tools` for Anthropic, carrying the same strict schema `Provider::Openai` sends. Reading it against the plain LLM probe is what makes the verdict sound rather than a guess at 404s: plain chat passing while the same request with tools fails means the model has no function calling; a response that carries no tool call means the endpoint took the tools but the model ignored them; both failing, or a timeout, stays an ordinary probe failure. The AI status card gains a Function calling (tools) row, an alert naming the fix for each of the two bad outcomes, and a failure reason. The hosting settings model field now says the assistant needs a tools-capable model and links super admins to the check. Refs #830 Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01DWCmRQ1ry26JnhA79s1ZKH * fix(ai-health): only call a refusal a refusal, and probe the route chat takes Two findings from review of the function-calling probe. A tools request can fail for reasons that say nothing about tool support: a 429, a 500, a dropped connection, an unreadable body. Reading any non-timeout failure as `:unsupported` sent the operator hunting for a new model over a transient blip. Only a 4xx the service chose to answer with — excluding the ones that mean "not now" or "not you" — is a refusal of the tools payload; everything else stays an ordinary probe failure. The bare 404 from OpenRouter that this page exists to explain still reads as missing function calling. `Provider::Openai#supports_responses_endpoint?` is the real routing decision and `OPENAI_SUPPORTS_RESPONSES_ENDPOINT` can flip it either way, so choosing the API from "is the endpoint custom" could probe Chat Completions while chat uses Responses, or the reverse — reporting on a path the assistant never takes. Ask the provider instead, and cache the two routes under separate keys. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01DWCmRQ1ry26JnhA79s1ZKH * fix(ai-health): confirm a tools refusal before blaming the model A client error on the tools request can mean "your tools payload" or "your request, tools or not" — an invalid schema, a route the endpoint does not serve, a model it will not run. Splitting those on the status code alone still put a 422 from an endpoint contract on the model's account and told the operator to go find another one. The probe now confirms it: when the tools request comes back a client error, it asks again with the tools taken off. Only if that lands is the tools payload what was turned down, and the probe says so with its own failure code — provider-agnostic, and no reading of error text for the word "tool", which would only ever fit the provider it was written against. Statuses that mean "not now" or "not you" (401, 402, 403, 408, 429) never get a second ask. `AiHealth` now just reports the probe's verdict instead of inferring one from the status. The troubleshooting fix no longer points at an OpenRouter free tier: those providers commonly log prompts and completions for training, and every assistant tool call carries accounts, transactions, and holdings. It points at the model recommendations already in this doc, and says why free tiers are the wrong place to look. Co-Authored-By: Claude Claude-Session: https://claude.ai/code/session_01DWCmRQ1ry26JnhA79s1ZKH --------- Co-authored-by: Claude --- app/models/ai_health.rb | 31 ++- app/models/ai_health/probe.rb | 159 +++++++++++++++ .../admin/system_health/_ai_status.html.erb | 27 +++ .../hostings/_openai_settings.html.erb | 6 + .../locales/views/admin/system_health/en.yml | 19 +- config/locales/views/settings/hostings/en.yml | 2 + docs/hosting/ai.md | 27 +++ .../admin/system_health_controller_test.rb | 44 ++++- test/models/ai_health/probe_test.rb | 184 ++++++++++++++++++ test/models/ai_health_test.rb | 97 +++++++++ test/system/admin/system_health_test.rb | 2 + 11 files changed, 593 insertions(+), 5 deletions(-) diff --git a/app/models/ai_health.rb b/app/models/ai_health.rb index e20fee294..bd39fddf2 100644 --- a/app/models/ai_health.rb +++ b/app/models/ai_health.rb @@ -19,8 +19,8 @@ class AiHealth :embedding_endpoint, :embedding_model, :embedding_dimensions, :pgvector_extension_available, :pgvector_extension_enabled, :pgvector_table_available, :qdrant_endpoint, :llm_probe, - :pdf_text_extraction_probe, :pdf_vision_processing_probe, - :vector_store_probe, :embedding_probe + :function_calling_probe, :pdf_text_extraction_probe, + :pdf_vision_processing_probe, :vector_store_probe, :embedding_probe def initialize(run_probes: true, force_probes: false) @run_probes = run_probes @@ -48,6 +48,20 @@ class AiHealth llm_probe.status end + # The assistant only answers through function calls, so an endpoint that + # serves plain chat but rejects the `tools` parameter still cannot power it. + # The probe confirms which of the two happened before reporting, so a model + # is only ever blamed for a refusal the service actually made. + def function_calling_status + return :unavailable unless llm_configured? + return :not_checked unless run_probes? + return :supported if function_calling_probe.passing? + return :not_used if function_calling_probe.failure_code == :no_tool_call + return :unsupported if function_calling_probe.failure_code == :tools_refused + + :failing + end + def llm_fallback? @effective_llm_protocol.present? && @effective_llm_protocol != @selected_llm_protocol end @@ -93,7 +107,8 @@ class AiHealth end def last_checked_at - [ llm_probe, pdf_text_extraction_probe, pdf_vision_processing_probe, vector_store_probe, embedding_probe ] + [ llm_probe, function_calling_probe, pdf_text_extraction_probe, pdf_vision_processing_probe, vector_store_probe, + embedding_probe ] .filter_map(&:checked_at) .max end @@ -129,6 +144,8 @@ class AiHealth @llm_model = effective_model(provider_for_details) @llm_endpoint = endpoint(provider_for_details) @llm_request_timeout = request_timeout(provider_for_details) + @openai_uses_responses_endpoint = @effective_llm_protocol == :openai && + safely(false) { @llm_provider.supports_responses_endpoint? } @pdf_processing_capable = safely(false) do @llm_provider&.supports_pdf_processing?(model: llm_model) end @@ -159,6 +176,7 @@ class AiHealth def load_probes @llm_probe = llm_configured? ? Probe.not_checked : Probe.not_configured + @function_calling_probe = llm_configured? ? Probe.not_checked : Probe.not_configured @pdf_text_extraction_probe = llm_configured? && @pdf_text_extraction_capable ? Probe.not_checked : Probe.not_configured @pdf_vision_processing_probe = llm_configured? && @pdf_vision_processing_capable ? Probe.not_checked : Probe.not_configured @vector_store_probe = vector_store_adapter.present? ? Probe.not_checked : Probe.not_configured @@ -174,6 +192,13 @@ class AiHealth model: llm_model, openai_compatible: @effective_llm_protocol == :openai && openai_compatible_endpoint? ) + @function_calling_probe = probe.function_calling( + provider: @effective_llm_protocol, + endpoint: @llm_raw_endpoint, + access_token: @llm_access_token, + model: llm_model, + use_responses_endpoint: @openai_uses_responses_endpoint + ) if @pdf_text_extraction_capable @pdf_text_extraction_probe = probe.pdf_text_extraction( provider: @effective_llm_protocol, diff --git a/app/models/ai_health/probe.rb b/app/models/ai_health/probe.rb index 21cb5798d..374df44a1 100644 --- a/app/models/ai_health/probe.rb +++ b/app/models/ai_health/probe.rb @@ -14,6 +14,24 @@ class AiHealth DEFAULT_TIMEOUT = 5 EMBEDDING_TEST_INPUT = "Sure AI health check" CHAT_TEST_INPUT = "Reply with OK." + FUNCTION_CALL_TEST_INPUT = "Call the sure_health_check tool with status set to ok." + FUNCTION_CALL_TEST_TOOL = { + name: "sure_health_check", + description: "Records the result of a Sure health check. Always call this tool.", + schema: { + type: "object", + properties: { + status: { type: "string", description: "Always the literal string \"ok\"." } + }, + required: [ "status" ], + additionalProperties: false + } + }.freeze + FUNCTION_CALL_MAX_RESPONSE_TOKENS = 256 + # 4xx statuses that mean "not now" or "not you" rather than "not this + # request": a retry, a payment, or a fixed credential clears them, so they + # are never worth re-asking without the tools. + TRANSIENT_HTTP_STATUSES = [ 401, 402, 403, 408, 429 ].freeze PDF_TEST_INSTITUTION = "SUREHEALTHCHECKBANK" PDF_TEST_LINES = [ "Bank Statement", @@ -89,6 +107,39 @@ class AiHealth end end + # The assistant reads financial data exclusively through function calls, so + # a model that rejects or ignores the `tools` parameter cannot power chat + # even when the plain check above passes. Providers report this + # inconsistently — OpenRouter answers a bare 404 — so ask the configured + # model for one trivial tool call and report what came back. The caller + # picks the OpenAI API, because `Provider::Openai#supports_responses_endpoint?` + # can be overridden: probing the route chat does not take proves nothing. + def function_calling(provider:, endpoint:, access_token:, model:, use_responses_endpoint: false) + run( + component: "function_calling", + provider_key: provider, + endpoint: endpoint, + model: model, + credential: access_token, + verification: use_responses_endpoint ? :responses_tool_call : :chat_tool_call + ) do + tool_called = case provider + when :openai + if use_responses_endpoint + openai_responses_tool_call?(access_token:, endpoint:, model:) + else + openai_chat_tool_call?(access_token:, endpoint:, model:) + end + when :anthropic + anthropic_tool_call?(access_token:, endpoint:, model:) + else + raise Failure, :unsupported_provider + end + + raise Failure, :no_tool_call unless tool_called + end + end + def pdf_text_extraction(provider:, endpoint:, access_token:, model:, openai_compatible: false) pdf_processing( provider: provider, @@ -264,6 +315,114 @@ class AiHealth response.is_a?(Hash) && response["choices"].is_a?(Array) && response["choices"].any? end + def openai_chat_tool_call?(access_token:, endpoint:, model:) + client = openai_client(access_token:, endpoint:) + parameters = { + model: model, + messages: [ { role: "user", content: FUNCTION_CALL_TEST_INPUT } ] + } + + response = confirming_tools_refusal( + tools: -> { client.chat(parameters: parameters.merge(tools: [ { type: "function", function: openai_test_tool } ])) }, + control: -> { client.chat(parameters: parameters) } + ) + + raise Failure, :invalid_response unless response.is_a?(Hash) && response["choices"].is_a?(Array) + + response.dig("choices", 0, "message", "tool_calls").present? + end + + def openai_responses_tool_call?(access_token:, endpoint:, model:) + client = openai_client(access_token:, endpoint:) + parameters = { + model: model, + input: [ { role: "user", content: FUNCTION_CALL_TEST_INPUT } ] + } + + response = confirming_tools_refusal( + tools: -> { client.responses.create(parameters: parameters.merge(tools: [ { type: "function" }.merge(openai_test_tool) ])) }, + control: -> { client.responses.create(parameters: parameters) } + ) + + raise Failure, :invalid_response unless response.is_a?(Hash) && response["output"].is_a?(Array) + + response["output"].any? { |item| item["type"] == "function_call" } + end + + def anthropic_tool_call?(access_token:, endpoint:, model:) + client = anthropic_client(access_token:, endpoint:) + parameters = { + model: model, + max_tokens: FUNCTION_CALL_MAX_RESPONSE_TOKENS, + messages: [ { role: "user", content: FUNCTION_CALL_TEST_INPUT } ] + } + + message = confirming_tools_refusal( + tools: -> { client.messages.create(**parameters, tools: [ anthropic_test_tool ]) }, + control: -> { client.messages.create(**parameters) } + ) + + Array(message.content).any? { |block| block_type(block) == "tool_use" } + end + + # A client error can mean "your tools payload" or "your request, tools or + # not" — an invalid schema, a route the endpoint does not serve, a model + # it will not run. Asking again without the tools is the only + # provider-agnostic way to tell those apart: if the same request lands + # once the tools come off, the tools are what was turned down. Reading + # the error text for the word "tool" instead would only ever fit the + # provider whose wording it was written against. + def confirming_tools_refusal(tools:, control:) + tools.call + rescue StandardError => error + raise error unless client_error?(error) + + begin + control.call + rescue StandardError + raise error + end + + raise Failure, :tools_refused + end + + def client_error?(error) + status = http_status(error).to_i + + status.between?(400, 499) && !status.in?(TRANSIENT_HTTP_STATUSES) + end + + # Mirrors the tool payload `Provider::Openai` sends for assistant + # functions, `strict` included, so an endpoint that only chokes on strict + # schemas is caught here rather than in chat. + def openai_test_tool + { + name: FUNCTION_CALL_TEST_TOOL[:name], + description: FUNCTION_CALL_TEST_TOOL[:description], + parameters: FUNCTION_CALL_TEST_TOOL[:schema], + strict: true + } + end + + # Anthropic names the schema differently and rejects OpenAI's `strict`. + def anthropic_test_tool + { + name: FUNCTION_CALL_TEST_TOOL[:name], + description: FUNCTION_CALL_TEST_TOOL[:description], + input_schema: FUNCTION_CALL_TEST_TOOL[:schema] + } + end + + def block_type(block) + raw = if block.respond_to?(:type) + block.type + elsif block.is_a?(Hash) + block[:type] || block["type"] + end + + raw.to_s + end + def valid_pdf_result?(result) result.is_a?(Provider::LlmConcept::PdfProcessingResult) && result.document_type == "bank_statement" && diff --git a/app/views/admin/system_health/_ai_status.html.erb b/app/views/admin/system_health/_ai_status.html.erb index 84ceb5608..b95317262 100644 --- a/app/views/admin/system_health/_ai_status.html.erb +++ b/app/views/admin/system_health/_ai_status.html.erb @@ -1,3 +1,10 @@ +<% function_calling_status = ai_health.function_calling_status %> +<% function_calling_class = case function_calling_status + when :supported then "text-success" + when :unsupported, :failing then "text-destructive" + else "text-warning" + end %> + <% pdf_checks = [ { key: :pdf_text_extraction, @@ -40,6 +47,14 @@ ) %> <% end %> + <% if function_calling_status.in?([ :unsupported, :not_used ]) %> + <%= render DS::Alert.new( + title: t("admin.system_health.show.ai.alerts.function_calling_#{function_calling_status}.title"), + message: t("admin.system_health.show.ai.alerts.function_calling_#{function_calling_status}.message"), + variant: function_calling_status == :unsupported ? :error : :warning + ) %> + <% end %> + <% pdf_checks.select { |check| check[:status] == :failing }.each do |check| %> <%= render DS::Alert.new( title: t("admin.system_health.show.ai.alerts.#{check[:key]}_failed.title"), @@ -123,6 +138,12 @@
<%= t("admin.system_health.show.ai.labels.endpoint") %>
<%= ai_health.llm_endpoint %>
+
+
<%= t("admin.system_health.show.ai.labels.function_calling") %>
+
+ <%= t("admin.system_health.show.ai.function_calling_statuses.#{function_calling_status}") %> +
+
<% pdf_checks.each do |check| %>
<%= t("admin.system_health.show.ai.labels.#{check[:key]}") %>
@@ -155,6 +176,12 @@
<%= t("admin.system_health.show.ai.failure_codes.#{ai_health.llm_probe.failure_code}") %>
<% end %> + <% if ai_health.function_calling_probe.failure_code %> +
+
<%= t("admin.system_health.show.ai.labels.function_calling_failure_reason") %>
+
<%= t("admin.system_health.show.ai.failure_codes.#{ai_health.function_calling_probe.failure_code}") %>
+
+ <% end %> <% pdf_checks.select { |check| check[:probe].failure_code }.each do |check| %>
<%= t("admin.system_health.show.ai.labels.#{check[:key]}_failure_reason") %>
diff --git a/app/views/settings/hostings/_openai_settings.html.erb b/app/views/settings/hostings/_openai_settings.html.erb index 7fbdc78dc..f9dac08ba 100644 --- a/app/views/settings/hostings/_openai_settings.html.erb +++ b/app/views/settings/hostings/_openai_settings.html.erb @@ -48,6 +48,12 @@ inputmode: "text", disabled: ENV["OPENAI_MODEL"].present?, data: { "auto-submit-form-target": "auto" } %> +

+ <%= t(".model_function_calling_help") %> + <% if Current.user&.super_admin? %> + <%= link_to t(".model_function_calling_link"), admin_system_health_path(tab: "ai"), class: "underline" %> + <% end %> +

<%= form.select :openai_json_mode, options_for_select( diff --git a/config/locales/views/admin/system_health/en.yml b/config/locales/views/admin/system_health/en.yml index 939e97c69..d1d52faab 100644 --- a/config/locales/views/admin/system_health/en.yml +++ b/config/locales/views/admin/system_health/en.yml @@ -42,6 +42,12 @@ en: custom_openai: title: The hosted vector-store check failed at this custom endpoint message: The configured endpoint did not pass the /v1/vector_stores liveness check. Most local and OpenAI-compatible endpoints do not implement that API. Chat and PDF processing may still work; use pgvector with a separate embeddings endpoint for the usual local setup. The failure was recorded in Settings → Debug logs and Rails.logger. + function_calling_unsupported: + title: The model does not support function calling + message: The endpoint served a plain chat request but rejected the same request carrying the tools parameter. The AI assistant reads accounts, transactions, and holdings through function calls, so chat keeps failing — often with an unhelpful 404 — until the model is one that supports tools. Pick a model whose provider documents function-calling support. The failure was recorded in Settings → Debug logs and Rails.logger. + function_calling_not_used: + title: The model answered without calling the tool it was asked to call + message: The endpoint accepted the tools parameter, but the model replied with text instead of calling the health-check tool. The AI assistant cannot read financial data without tool calls, so chat may answer with nothing or with invented figures. Prefer a model whose provider documents function-calling support. llm_probe_failed: title: The LLM live check failed message: Sure could not verify both the configured endpoint and model. The failure was recorded in Settings → Debug logs and Rails.logger. @@ -65,7 +71,7 @@ en: message: Check the adapter requirements below. For pgvector, the database must provide the vector extension or an existing vector_store_chunks table. llm: title: LLM and PDF processing - description: The live checks verify the configured model, then separately exercise text extraction and vision/native processing with a one-page synthetic PDF containing no customer data. Credential values are never displayed. + description: The live checks verify the configured model, ask it for one trivial function call the way the assistant does, then separately exercise text extraction and vision/native processing with a one-page synthetic PDF containing no customer data. Credential values are never displayed. vector_store: title: Vector store and document search description: The live check queries the hosted vector-store API, or tests the pgvector table and creates one short embedding without storing it. @@ -74,6 +80,7 @@ en: effective_provider: Effective provider model: Model endpoint: Endpoint + function_calling: Function calling (tools) pdf_text_extraction: PDF text-extraction path pdf_vision_processing: PDF vision/native path request_timeout: Request timeout @@ -87,6 +94,7 @@ en: embedding_model: Embedding model embedding_endpoint: Embedding endpoint embedding_dimensions: Embedding dimensions + function_calling_failure_reason: Function-calling failure reason pdf_text_extraction_failure_reason: Text-extraction failure reason pdf_vision_processing_failure_reason: Vision/native failure reason storage_probe: pgvector storage check @@ -118,6 +126,13 @@ en: not_available: Not available available: Available not_found: Not found + function_calling_statuses: + supported: Live tool call succeeded + unsupported: Not supported by the effective provider/model + not_used: Tools accepted, but the model called none + failing: Live check failed + not_checked: Not checked + unavailable: Unavailable until an LLM provider is configured pdf_statuses: supported: Synthetic PDF check passed failing: Synthetic PDF check failed @@ -138,6 +153,8 @@ en: missing: Not configured failure_codes: model_not_available: The configured model was not returned by the provider + no_tool_call: The model answered without calling the health-check tool + tools_refused: The service served the same request without tools and refused it with them invalid_response: The service returned an unexpected response dimensions_mismatch: The embedding vector dimensions do not match the configured dimensions extension_not_enabled: The PostgreSQL vector extension is not enabled diff --git a/config/locales/views/settings/hostings/en.yml b/config/locales/views/settings/hostings/en.yml index 1b2068598..23af61131 100644 --- a/config/locales/views/settings/hostings/en.yml +++ b/config/locales/views/settings/hostings/en.yml @@ -127,6 +127,8 @@ en: uri_base_placeholder: "https://api.openai.com/v1 (default)" model_label: Model (Optional) model_placeholder: "gpt-4.1 (default)" + model_function_calling_help: "The AI assistant reads your data through function calls, so pick a model your provider documents as supporting tools/function calling. Models without it fail chat with an opaque provider error." + model_function_calling_link: Check the configured model json_mode_label: JSON Mode json_mode_auto: Auto (recommended) json_mode_strict: Strict (best for thinking models) diff --git a/docs/hosting/ai.md b/docs/hosting/ai.md index a8157ec3e..f32c44992 100644 --- a/docs/hosting/ai.md +++ b/docs/hosting/ai.md @@ -1082,6 +1082,30 @@ ollama list # See what's installed ollama pull model-name # Install a model ``` +### Chat Fails With a Bare "404" (Model Without Function Calling) + +**Symptom:** The assistant answers every message with an unexplained `404` (or +another opaque provider error), while the same endpoint and model work for +auto-categorization. + +**Cause:** The assistant reads accounts, transactions, and holdings through +function calls, so every chat request carries a `tools` payload. A model +without function-calling support rejects it — OpenRouter answers `404` for +models such as `tngtech/deepseek-r1t2-chimera:free`. + +**Confirm it:** Open **System health → AI status** +(`/admin/system_health?tab=ai`). **Function calling (tools)** reports *Not +supported by the effective provider/model* when the endpoint served the plain +chat check but rejected the same request carrying tools, and *Tools accepted, +but the model called none* when the model answered with text instead of +calling the tool. + +**Fix:** Set `OPENAI_MODEL` to a model your provider documents as supporting +tools/function calling — see [For Chat Assistant](#for-chat-assistant) above — +then run the checks again. Free tiers are a poor place to look: providers +commonly log their prompts and completions for training, and the assistant +sends your accounts, transactions, and holdings in every tool call. + ### "Fixed prompt tokens exceed context budget" **Symptom:** Auto-categorization or merchant detection fails immediately with an error like: @@ -1408,6 +1432,9 @@ live checks against the effective configuration: - OpenAI-compatible and Anthropic providers must return the configured model from their models API. +- The configured model must complete one trivial function call, sent the way + the assistant sends its own tools. A model that serves plain chat but rejects + or ignores the `tools` parameter cannot answer questions about your data. - The hosted OpenAI vector-store adapter must answer a list request without creating or changing a store. - The pgvector adapter must have its extension enabled, its chunks table diff --git a/test/controllers/admin/system_health_controller_test.rb b/test/controllers/admin/system_health_controller_test.rb index 5af9070db..2d6f81fa1 100644 --- a/test/controllers/admin/system_health_controller_test.rb +++ b/test/controllers/admin/system_health_controller_test.rb @@ -3,7 +3,8 @@ require "test_helper" class Admin::SystemHealthControllerTest < ActionDispatch::IntegrationTest AI_ENVIRONMENT = %w[ OPENAI_ACCESS_TOKEN OPENAI_URI_BASE OPENAI_MODEL OPENAI_REQUEST_TIMEOUT - OPENAI_SUPPORTS_PDF_PROCESSING ANTHROPIC_ACCESS_TOKEN ANTHROPIC_API_KEY + OPENAI_SUPPORTS_PDF_PROCESSING OPENAI_SUPPORTS_RESPONSES_ENDPOINT + ANTHROPIC_ACCESS_TOKEN ANTHROPIC_API_KEY ANTHROPIC_BASE_URL ANTHROPIC_MODEL ANTHROPIC_REQUEST_TIMEOUT VECTOR_STORE_PROVIDER EMBEDDING_URI_BASE EMBEDDING_MODEL EMBEDDING_DIMENSIONS EMBEDDING_ACCESS_TOKEN QDRANT_URL QDRANT_API_KEY @@ -19,6 +20,7 @@ class Admin::SystemHealthControllerTest < ActionDispatch::IntegrationTest Setting.stubs(:anthropic_base_url).returns(nil) Setting.stubs(:anthropic_model).returns(nil) AiHealth::Probe.any_instance.stubs(:llm).returns(probe_result(:passing)) + AiHealth::Probe.any_instance.stubs(:function_calling).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:pdf_text_extraction).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:pdf_vision_processing).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:openai_vector_store).returns(probe_result(:passing)) @@ -109,6 +111,7 @@ class Admin::SystemHealthControllerTest < ActionDispatch::IntegrationTest sign_in users(:sure_support_staff) stub_healthy_sidekiq AiHealth::Probe.any_instance.expects(:llm).never + AiHealth::Probe.any_instance.expects(:function_calling).never AiHealth::Probe.any_instance.expects(:pdf_text_extraction).never AiHealth::Probe.any_instance.expects(:pdf_vision_processing).never AiHealth::Probe.any_instance.expects(:openai_vector_store).never @@ -156,6 +159,45 @@ class Admin::SystemHealthControllerTest < ActionDispatch::IntegrationTest assert_no_match(/local-token|uri-secret|query-secret/, response.body) end + test "AI status names the missing function-calling support behind an unhelpful chat error" do + sign_in users(:sure_support_staff) + stub_healthy_sidekiq + AiHealth::Probe.any_instance.stubs(:function_calling).returns( + probe_result(:failing, failure_code: :tools_refused, http_status: 404) + ) + + with_ai_environment( + "OPENAI_ACCESS_TOKEN" => "router-secret", + "OPENAI_URI_BASE" => "https://openrouter.ai/api/v1", + "OPENAI_MODEL" => "tngtech/deepseek-r1t2-chimera:free" + ) do + get admin_system_health_url(tab: "ai") + end + + assert_response :success + assert_select "[data-testid='function-calling-status']", text: /Not supported by the effective provider/ + assert_match(/The model does not support function calling/, response.body) + assert_match(/Function-calling failure reason/, response.body) + assert_no_match(/router-secret/, response.body) + end + + test "AI status separates a model that ignores tools from one that cannot use them" do + sign_in users(:sure_support_staff) + stub_healthy_sidekiq + AiHealth::Probe.any_instance.stubs(:function_calling).returns( + probe_result(:failing, failure_code: :no_tool_call) + ) + + with_ai_environment("OPENAI_ACCESS_TOKEN" => "sk-secret-openai") do + get admin_system_health_url(tab: "ai") + end + + assert_response :success + assert_select "[data-testid='function-calling-status']", text: /Tools accepted, but the model called none/ + assert_match(/answered without calling the tool it was asked to call/, response.body) + assert_no_match(/The model does not support function calling/, response.body) + end + test "AI status reports text and vision PDF probes separately" do sign_in users(:sure_support_staff) stub_healthy_sidekiq diff --git a/test/models/ai_health/probe_test.rb b/test/models/ai_health/probe_test.rb index 5e7430cd0..7b5e28a0f 100644 --- a/test/models/ai_health/probe_test.rb +++ b/test/models/ai_health/probe_test.rb @@ -378,4 +378,188 @@ class AiHealth::ProbeTest < ActiveSupport::TestCase assert result.failing? assert_equal :dimensions_mismatch, result.failure_code end + + test "function-calling probe sends the assistant's tools payload and passes when the model calls one" do + endpoint = "https://openrouter.example.test/api/v1" + request = stub_request(:post, "#{endpoint}/chat/completions") + .with { |req| + body = JSON.parse(req.body) + tool = body.dig("tools", 0, "function") + + body["model"] == "tools-model" && + body.dig("messages", 0, "content") == AiHealth::Probe::FUNCTION_CALL_TEST_INPUT && + tool["name"] == AiHealth::Probe::FUNCTION_CALL_TEST_TOOL[:name] && + tool["strict"] == true + } + .to_return( + status: 200, + headers: { "Content-Type" => "application/json" }, + body: { + choices: [ + { + message: { + tool_calls: [ + { id: "call_1", type: "function", function: { name: "sure_health_check", arguments: "{}" } } + ] + } + } + ] + }.to_json + ) + + result = @probe.function_calling( + provider: :openai, + endpoint: endpoint, + access_token: "token", + model: "tools-model" + ) + + assert result.passing? + assert_requested request + end + + test "function-calling probe fails when the model answers without calling the tool" do + endpoint = "https://openrouter.example.test/api/v1" + stub_request(:post, "#{endpoint}/chat/completions").to_return( + status: 200, + headers: { "Content-Type" => "application/json" }, + body: { choices: [ { message: { content: "ok" } } ] }.to_json + ) + Rails.logger.stubs(:error) + DebugLogEntry.stubs(:capture) + + result = @probe.function_calling( + provider: :openai, + endpoint: endpoint, + access_token: "token", + model: "chat-only-model" + ) + + assert result.failing? + assert_equal :no_tool_call, result.failure_code + end + + test "function-calling probe confirms a refusal by re-asking without the tools" do + endpoint = "https://openrouter.example.test/api/v1" + stub_tools_request(endpoint).to_return( + status: 404, + headers: { "Content-Type" => "application/json" }, + body: { error: { message: "No endpoints found that support tool use." } }.to_json + ) + control = stub_control_request(endpoint).to_return( + status: 200, + headers: { "Content-Type" => "application/json" }, + body: { choices: [ { message: { content: "ok" } } ] }.to_json + ) + Rails.logger.stubs(:error) + DebugLogEntry.stubs(:capture) + + result = @probe.function_calling( + provider: :openai, + endpoint: endpoint, + access_token: "token", + model: "tngtech/deepseek-r1t2-chimera:free" + ) + + assert result.failing? + assert_equal :tools_refused, result.failure_code + assert_requested control + end + + test "function-calling probe does not blame the tools for a client error the request gets either way" do + endpoint = "https://models.example.test/v1" + stub_tools_request(endpoint).to_return(status: 422, body: "{}", headers: { "Content-Type" => "application/json" }) + control = stub_control_request(endpoint).to_return( + status: 422, + body: "{}", + headers: { "Content-Type" => "application/json" } + ) + Rails.logger.stubs(:error) + DebugLogEntry.stubs(:capture) + + result = @probe.function_calling( + provider: :openai, + endpoint: endpoint, + access_token: "token", + model: "picky-model" + ) + + assert result.failing? + assert_equal :request_failed, result.failure_code + assert_equal 422, result.http_status + assert_requested control + end + + test "function-calling probe does not re-ask when the service said not now" do + endpoint = "https://models.example.test/v1" + stub_tools_request(endpoint).to_return(status: 429, body: "{}", headers: { "Content-Type" => "application/json" }) + control = stub_control_request(endpoint) + Rails.logger.stubs(:error) + DebugLogEntry.stubs(:capture) + + result = @probe.function_calling( + provider: :openai, + endpoint: endpoint, + access_token: "token", + model: "busy-model" + ) + + assert result.failing? + assert_equal :request_failed, result.failure_code + assert_equal 429, result.http_status + assert_not_requested control + end + + test "function-calling probe uses the responses endpoint when the assistant would" do + request = stub_request(:post, "https://api.openai.example.test/v1/responses") + .with { |req| + tool = JSON.parse(req.body).dig("tools", 0) + + tool["type"] == "function" && tool["name"] == AiHealth::Probe::FUNCTION_CALL_TEST_TOOL[:name] + } + .to_return( + status: 200, + headers: { "Content-Type" => "application/json" }, + body: { output: [ { type: "function_call", name: "sure_health_check", arguments: "{}" } ] }.to_json + ) + + result = @probe.function_calling( + provider: :openai, + endpoint: "https://api.openai.example.test/v1", + access_token: "token", + model: "gpt-4.1", + use_responses_endpoint: true + ) + + assert result.passing? + assert_requested request + end + + test "function-calling probe reads an Anthropic tool_use block" do + response = Struct.new(:content).new([ Struct.new(:type, :name).new(:tool_use, "sure_health_check") ]) + messages = mock("anthropic_messages") + messages.expects(:create).with do |params| + params[:tools].first[:input_schema] == AiHealth::Probe::FUNCTION_CALL_TEST_TOOL[:schema] && + !params[:tools].first.key?(:strict) + end.returns(response) + @probe.stubs(:anthropic_client).returns(stub(messages: messages)) + + result = @probe.function_calling( + provider: :anthropic, + endpoint: "https://api.anthropic.com", + access_token: "token", + model: "claude-sonnet-4-6" + ) + + assert result.passing? + end + + private + def stub_tools_request(endpoint) + stub_request(:post, "#{endpoint}/chat/completions").with { |request| JSON.parse(request.body).key?("tools") } + end + + def stub_control_request(endpoint) + stub_request(:post, "#{endpoint}/chat/completions").with { |request| !JSON.parse(request.body).key?("tools") } + end end diff --git a/test/models/ai_health_test.rb b/test/models/ai_health_test.rb index 04763ff36..3c3b87235 100644 --- a/test/models/ai_health_test.rb +++ b/test/models/ai_health_test.rb @@ -4,6 +4,7 @@ class AiHealthTest < ActiveSupport::TestCase AI_ENVIRONMENT = %w[ OPENAI_ACCESS_TOKEN OPENAI_URI_BASE OPENAI_MODEL ANTHROPIC_ACCESS_TOKEN ANTHROPIC_API_KEY VECTOR_STORE_PROVIDER + OPENAI_SUPPORTS_RESPONSES_ENDPOINT ].index_with(nil).freeze setup do @@ -42,7 +43,103 @@ class AiHealthTest < ActiveSupport::TestCase end end + test "a confirmed tools refusal reads as missing function-calling support" do + health = probed_health(llm: :passing, function_calling: failing_result(:tools_refused, http_status: 404)) + + assert_equal :unsupported, health.function_calling_status + end + + test "a service that fell over is not read as a model without function calling" do + [ + failing_result(:request_failed, http_status: 422), + failing_result(:request_failed, http_status: 429), + failing_result(:request_failed, http_status: 500), + failing_result(:request_failed), + failing_result(:invalid_response), + failing_result(:timeout) + ].each do |probe_result| + health = probed_health(llm: :passing, function_calling: probe_result) + + assert_equal :failing, health.function_calling_status, + "#{probe_result.failure_code} #{probe_result.http_status} should not blame the model" + end + end + + test "a model that answers without calling the tool is reported separately" do + health = probed_health(llm: :passing, function_calling: failing_result(:no_tool_call)) + + assert_equal :not_used, health.function_calling_status + end + + test "a refusal still reads as unsupported when the plain LLM check is failing too" do + health = probed_health(llm: :failing, function_calling: failing_result(:tools_refused)) + + assert_equal :unsupported, health.function_calling_status + end + + test "function calling is probed through the API the assistant would use" do + { + { "OPENAI_URI_BASE" => nil } => true, + { "OPENAI_URI_BASE" => "https://openrouter.ai/api/v1" } => false, + { "OPENAI_URI_BASE" => "https://openrouter.ai/api/v1", "OPENAI_SUPPORTS_RESPONSES_ENDPOINT" => "true" } => true, + { "OPENAI_SUPPORTS_RESPONSES_ENDPOINT" => "false" } => false + }.each do |environment, use_responses_endpoint| + AiHealth::Probe.any_instance.stubs(:llm).returns(result(:passing)) + AiHealth::Probe.any_instance.stubs(:pdf_text_extraction).returns(result(:passing)) + AiHealth::Probe.any_instance.stubs(:pdf_vision_processing).returns(result(:passing)) + AiHealth::Probe.any_instance.expects(:function_calling) + .with(has_entry(use_responses_endpoint: use_responses_endpoint)) + .returns(result(:passing)) + + ClimateControl.modify( + AI_ENVIRONMENT.merge( + "OPENAI_ACCESS_TOKEN" => "test-token", + "OPENAI_MODEL" => "test-model" + ).merge(environment) + ) { AiHealth.new } + end + end + + test "function calling is only checked alongside the other live probes" do + with_openai_endpoint("https://openrouter.ai/api/v1") do |health| + assert_equal :not_checked, health.function_calling_status + end + end + private + def probed_health(llm:, function_calling:) + AiHealth::Probe.any_instance.stubs(:llm).returns(result(llm)) + AiHealth::Probe.any_instance.stubs(:function_calling).returns(function_calling) + AiHealth::Probe.any_instance.stubs(:pdf_text_extraction).returns(result(:passing)) + AiHealth::Probe.any_instance.stubs(:pdf_vision_processing).returns(result(:passing)) + + ClimateControl.modify( + AI_ENVIRONMENT.merge( + "OPENAI_ACCESS_TOKEN" => "test-token", + "OPENAI_URI_BASE" => "https://openrouter.ai/api/v1", + "OPENAI_MODEL" => "test-model" + ) + ) { AiHealth.new } + end + + def result(status) + AiHealth::Probe::Result.new( + status: status, + checked_at: Time.current, + failure_code: nil, + http_status: nil + ) + end + + def failing_result(failure_code, http_status: nil) + AiHealth::Probe::Result.new( + status: :failing, + checked_at: Time.current, + failure_code: failure_code, + http_status: http_status + ) + end + def with_openai_endpoint(endpoint) ClimateControl.modify( AI_ENVIRONMENT.merge( diff --git a/test/system/admin/system_health_test.rb b/test/system/admin/system_health_test.rb index 4b102dd4b..cdd4c41cc 100644 --- a/test/system/admin/system_health_test.rb +++ b/test/system/admin/system_health_test.rb @@ -9,6 +9,7 @@ class Admin::SystemHealthTest < ApplicationSystemTestCase Setting.stubs(:openai_model).returns(nil) stub_healthy_sidekiq AiHealth::Probe.any_instance.stubs(:llm).returns(probe_result(:passing)) + AiHealth::Probe.any_instance.stubs(:function_calling).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:pdf_text_extraction).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:pdf_vision_processing).returns(probe_result(:passing)) AiHealth::Probe.any_instance.stubs(:openai_vector_store).returns(probe_result(:passing)) @@ -28,6 +29,7 @@ class Admin::SystemHealthTest < ApplicationSystemTestCase assert_current_path admin_system_health_path(tab: "ai") assert_selector "button[role='tab'][aria-selected='true']", text: "AI status" assert_text "Live check passed" + assert_text "Live tool call succeeded" assert_text "PDF text-extraction path" assert_text "PDF vision/native path" assert_text "Synthetic PDF check passed", count: 2