* Fix PDF vision path failing when poppler-utils is missing
The Docker image omitted poppler-utils, so the OpenAI PDF vision path
(which renders pages with pdftoppm before sending them upstream) always
failed and the admin AI status page surfaced a generic "request failed"
code rather than a useful reason.
- install poppler-utils in the Docker base image (pdftoppm for the
vision render path)
- add an optional failure_code to Provider::Error, so probe errors can
carry a machine-readable reason
- in Provider::Openai::PdfProcessor#convert_pdf_to_images, check
pdftoppm's return value and -- when the binary is genuinely missing,
raise a Provider::Openai::Error with failure_code :render_missing_binary
while preserving the existing [] fallback for other render failures
- register the :render_missing_binary code in the admin locale so the AI
status page shows a concrete, actionable message instead of "request
failed"
- AiHealth::Probe#failure_code now falls through to the default codes
when an error's failure_code is nil, instead of returning nil
- regression tests covering both the missing-binary and the
present-but-fails cases
Fixes the "PDF vision/native path" system check on instances where the
image is built from the checked-in Dockerfile.
* Fix missing-binary detection and preserve failure_code across the error boundary
- convert_pdf_to_images: Kernel#system returns nil (not false) when the executable is absent; raise the coded error on rendered.nil? so a missing pdftoppm yields :render_missing_binary instead of a blank conversion. - Provider::Error#as_json + default_error_transformer: carry failure_code through serialization and error re-wrapping. - Drop the binary_missing? helper (nil result is the authoritative signal) and update regression tests to stub the real nil return value. - Add coverage for failure_code serialization/transformation.
* Fix Provider::Error transformer syntax error and cover Faraday branch
Addresses jjmata's blocking change-request: app/models/provider.rb:54
used a postfix `if` modifier inside a hash-argument/method-call argument
list, which is not valid Ruby. `ruby -c` failed to parse the file, so the
base class every provider inherits from could not autoload and the whole
app (boot, requests, jobs, tests) was down.
Rewrite default_error_transformer to build the optional failure_code
keyword once (only when the error exposes a truthy code) and splat it, so
both the Faraday::Error branch and the generic branch carry the code with
no syntax error and no nil kwarg.
Also add the two Faraday-branch regression tests that were missing
(failure_code preservation + response-body-to-details extraction),
guarding this exact class of broken-argument bug with real assertions.
Verification: ruby -c passes on provider.rb, pdf_processor.rb, probe.rb,
and both test files; plus a behavioral harness against the real
provider.rb source covering the coded/plain/nil-code, Faraday-coded,
Faraday-no-code, nil-response, and generic-error paths (19/19 pass).
Ref: we-promise/sure#3275
* Document the methods added or changed by this PR
* Add poppler-utils to devcontainer image
---------
Co-authored-by: hermes-on-behalf-of-jon <hermes@nousresearch.com>
Co-authored-by: jaysbeekay <jaysbeekay@users.noreply.github.com>
Co-authored-by: sure-admin <sure-admin@splashblot.com>
* Add synthetic PDF health checks
* Require exact marker in PDF health probes
* Report PDF health paths separately
* Refactor application code
* Remove unrelated schema dump changes
* Simplify synthetic PDF validation
---------
Signed-off-by: Juan José Mata <juanjo.mata@gmail.com>
* fix(providers): capture API error response body in PDF processor span output
Anthropic and OpenAI PDF processing errors only logged the exception
message, dropping the parsed response body that usually explains the
failure. Add safe_error_body to both providers' UsageRecorder concerns
and include it in the langfuse span output on failure, with tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013SYp89xEYfTkrxw8HcCUQ8
* fix(providers): allowlist PDF processor error fields sent to Langfuse
safe_error_body forwarded the entire upstream error body into the
langfuse span output. For custom OpenAI-compatible providers/proxies
(and the analogous Anthropic path), that body can echo request
content from the financial document being processed. Replace it with
safe_error_detail, which extracts only type/message/code/request_id
instead of the raw body.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013SYp89xEYfTkrxw8HcCUQ8
* test(openai): cover request_id extraction in PDF processor error_detail
The safe_error_detail request_id path (error.response_headers) had no
test coverage. Stub response_headers with x-request-id and assert it
appears in error_detail.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013SYp89xEYfTkrxw8HcCUQ8
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>