Commit Graph
7 Commits
Author SHA1 Message Date
6439a731ab feat(self-hosting): surface Sidekiq-unhealthy nudge + admin system health page (#1906)
* feat(self-hosting): surface Sidekiq-unhealthy nudge + admin system health page (#1481)

When the Sidekiq worker container isn't running — the most common Docker
Compose misconfiguration in self-hosted setups — every background job
silently never executes. Balance calculations, net-worth updates, and
account syncs stall. The UI shows zeros and "No balance data available
for this date" without explaining why (#1481, #1047).

Per jjmata's resolution on the issue, this PR ships both halves of the
fix in one pass:

1. A user-facing nudge banner that appears on every authenticated page
   when Sidekiq isn't processing jobs. Tells the user their data may be
   stale; doesn't pretend zeros are real.

2. An admin-only deep link from that banner into a new
   `/settings/admin/system_health` page (super-admin gated, matching the
   existing admin namespace contract) showing live Sidekiq state:
   process count, last heartbeat, max queue latency, job counters, and
   per-queue depth.

## What changed

- New `SidekiqHealth` PORO (`app/models/sidekiq_health.rb`) eagerly
  loads ProcessSet + Queue + Stats in one pass and exposes `healthy?`
  plus a stable `reason` symbol (`:redis_unreachable`,
  `:no_worker_processes`, `:stale_heartbeat`, `:queue_backed_up`).
  Any Redis/Sidekiq failure during the eager load is caught and
  surfaced as `:redis_unreachable` so a degraded broker never crashes
  the layout.

- `ApplicationController#current_sidekiq_health` memoizes a single
  instance per request via `helper_method` so the layout, banner
  partial, and any controller checks share one Redis round-trip.

- New `app/views/shared/_sidekiq_health_banner.html.erb` rendered from
  `_htmldoc.html.erb` when `Current.user` is present and the health
  check is failing. Banner shows the user-facing message to everyone;
  the "View system health" CTA + reason detail are gated on
  `Current.user&.super_admin?`.

- New `Admin::SystemHealthController#show` (inherits the existing
  `Admin::BaseController`, so super-admin gating is enforced for free)
  + view rendering status, counters, and per-queue breakdown.

- Routes: `resource :system_health, only: :show` inside the existing
  `namespace :admin`.

- Settings nav: new "System health" entry under the Advanced section,
  gated on `super_admin?` to match `sso_providers_label` and
  `users_label`.

- i18n: new `shared.sidekiq_health_banner.*` keys (title, body, CTA,
  per-reason explanations) and a full `admin.system_health.show.*`
  namespace for the new admin page. English-only, matching how
  `ds.pill.*` and other DS keys are scoped.

## Why

- jjmata: "Let's take both approaches ... a nudge about 'data
  unavailable' which hyperlinks to the admin UI if you are an admin
  only (not for other types of users) sounds like the best path forward.
  **Any takers for the PR?**" (#1481)
- smurfpandey: "We can add a section in Settings for superadmins to see
  'health' of the application/host."
- The detection signal is conservative on purpose:
  - `PROCESS_HEARTBEAT_TIMEOUT = 2.minutes` tolerates deploy restarts
    and brief Redis blips without flapping.
  - `LATENCY_THRESHOLD = 5.minutes` is well above the sync-job tail
    under default `config/sidekiq.yml` concurrency.

## Validation

This worktree runs on Windows without a local Ruby toolchain, so I
could not run `bin/rubocop`, `bundle exec erb_lint`, `bin/brakeman`, or
`bin/rails test` locally. CI will run the full matrix on the PR:

- `lint` — `bin/rubocop -f github`
- `lint_js` — `npm run lint` (no JS touched, should be green)
- `scan_ruby` — `bin/brakeman --no-pager`
- `scan_js` — `bin/importmap audit`
- `test_unit` — `bin/rails test` (includes 7 new tests under
  `test/models/sidekiq_health_test.rb` and 4 new under
  `test/controllers/admin/system_health_controller_test.rb`)
- `test_system` — `DISABLE_PARALLELIZATION=true bin/rails test:system`
- `pipelock` — secret + agent-security diff scan

Manual checks done in this worktree:

- Re-read `CONTRIBUTING.md` and `.cursor/rules/project-conventions.mdc`.
  PORO under `app/models/` per Convention 2. No new gem dependency per
  Convention 1. Banner uses semantic tokens (`bg-warning/10`,
  `text-warning`) per the design-system rules. No `lucide_icon` direct
  call — uses the `icon` helper per CLAUDE.md.
- Confirmed `Sidekiq::ProcessSet` / `Sidekiq::Queue` / `Sidekiq::Stats`
  are the same APIs Sidekiq 7+ exposes (we're on Sidekiq 8.x per the
  `Gemfile.lock` comment in `config/initializers/sidekiq.rb`).
- Tests stub `Sidekiq::ProcessSet.new` / `Sidekiq::Queue.all` /
  `Sidekiq::Stats.new` so the suite doesn't need Redis populated.
- The admin route lives inside the existing `namespace :admin` so
  `Admin::BaseController#require_super_admin!` enforces auth — no new
  authorization surface added.

## Notes

- No public API endpoints, no rswag specs, no OpenAPI changes.
- No migrations, no model changes outside the new PORO.
- No background jobs touched.
- English-only locale entry, mirroring the `ds.*` / `admin.invitations.*`
  precedent in this repo. Other locales fall back to English.
- Detection thresholds are constants on `SidekiqHealth` so they're easy
  to tune from a follow-up PR if the defaults turn out to flap on any
  real-world deployment.
- The banner positions itself at `top-20` (below the impersonation /
  super-admin bars) and uses `z-40` (below the `z-50` notification
  tray). Single-screen overlap with mobile flash toasts is acceptable
  for V1.

Refs: #1481, #1047

* fix(self-hosting): address review on Sidekiq health PR (#1481)

- `Admin::SystemHealthController#show` now reads from the request-memoized
  `current_sidekiq_health` instead of building a fresh `SidekiqHealth.new`,
  so the controller and the layout banner share one Redis round-trip.
- `SidekiqHealth#reason` now treats `last_heartbeat_at.nil?` the same as a
  stale beat: a registered process that hasn't published a heartbeat is
  not "healthy". Previously the check short-circuited on the nil guard
  and silently fell through to the queue-latency branch. Added a unit
  test covering the `ProcessSet` entry with `"beat" => nil` case.
- Settings nav: switched the "System health" entry's icon from `activity`
  to `heart-pulse` so it no longer duplicates the LLM Usage icon.
- Routes: dropped the redundant `controller: "system_health"` option from
  the `resource :system_health` declaration — Rails infers
  `Admin::SystemHealthController` from the namespace, matching the style
  of the sibling `:sso_providers`, `:users`, `:invitations`, and
  `:families` admin resources.

* fix(self-hosting): scope + cache Sidekiq health, admin-only banner (#1481)

Addresses the second round of maintainer review on the Sidekiq health PR.

- Skip the check entirely in managed mode. `current_sidekiq_health`
  returns `nil` unless `Rails.application.config.app_mode.self_hosted?`,
  so authenticated requests in managed deployments add zero Redis
  round-trips for this feature.
- Cache the snapshot across requests via `SidekiqHealth.current`
  (Rails.cache, TTL `CACHE_TTL` = 60s default, env-overridable). The
  per-request memoization on `ApplicationController` is preserved on
  top, so even back-to-back self-hosted pages share one fetch.
- Make thresholds operator-tunable. `PROCESS_HEARTBEAT_TIMEOUT`,
  `LATENCY_THRESHOLD`, and the new `CACHE_TTL` read from
  `SIDEKIQ_HEALTH_HEARTBEAT_TIMEOUT`, `SIDEKIQ_HEALTH_LATENCY_THRESHOLD`,
  and `SIDEKIQ_HEALTH_CACHE_TTL` env vars (seconds), with the previous
  values as defaults. Comments now explain the tuning rationale.
- Gate the banner on `Current.user&.super_admin?` at the layout level
  rather than rendering a vague warning to family members who can't
  act on it. The partial no longer carries an internal admin check
  since the call site does it; non-admins see nothing.
- Replace the hard-coded `top-20` offset with a computed offset based
  on which impersonation bars are visible (`top-4` / `top-20` / `top-36`)
  so the banner doesn't collide with the super-admin or approval bars
  when both are stacked above it.
- `Admin::SystemHealthController#show` now bypasses the cache
  (`SidekiqHealth.expire_cache!` + `SidekiqHealth.new`) so an operator
  who just restarted the worker sees fresh state instead of a stale
  60-second snapshot. Also lets the page render in managed mode where
  `current_sidekiq_health` is nil.
- Tests: add coverage for `.current` cache reuse and `.expire_cache!`
  forcing a re-query, swapping `Rails.cache` to a MemoryStore since the
  test env defaults to `:null_store`.

* fix(self-hosting): route singular resource + drop assert_same on cached snapshot (#1481)

Two CI failures surfaced once the full pipeline ran on this branch for
the first time (it was gated on contributor approval until d04b78e):

- Admin system-health controller tests returned 404. Singular
  `resource :system_health` in `config/routes.rb` makes Rails infer
  `Admin::SystemHealthsController` (it pluralizes the controller name
  even for singular resources), but the controller file is named
  `system_health_controller.rb` / `Admin::SystemHealthController`.
  Restore the explicit `controller: "system_health"` override that
  the previous "address review" commit dropped on the (mistaken)
  premise that Rails would infer it from the namespace — the sibling
  admin routes all use plural `resources` so they round-trip cleanly,
  this one doesn't. Comment now spells the gotcha out so the next
  reviewer doesn't try to "simplify" it again.
- `SidekiqHealthTest#test_current_memoizes_across_calls_inside_the_cache_TTL`
  used `assert_same` on the two returns from `SidekiqHealth.current`.
  `ActiveSupport::Cache::MemoryStore` defaults to `dup_values: true`
  and Marshals on read, so a cache hit returns an `==`-equal but
  `equal?`-different instance. Replace the identity check with the
  behavioral assertion we actually care about: re-stub `ProcessSet`
  to raise on the second call, then assert the second `current`
  return is still healthy (proving Redis was not re-queried).

* fix(i18n): drop redundant inline default on system_health nav label (#1481)

`system_health_label` is already defined in
config/locales/views/settings/en.yml, so the inline
`default: "System health"` was a hard-coded English string in the
template (DS Drift Patrol Rule 5). Use the bare locale lookup like the
sibling nav entries.

---------

Co-authored-by: John Baillie <johnbaillie2007@gmail.com>
Co-authored-by: Khaostica <256858950+Khaostica@users.noreply.github.com>
2026-08-22 07:17:29 +02:00
Juan José Mata 172301f875 Fix SSO provider settings updates (#2210) 2026-06-06 09:14:50 +02:00
Juan José MataandClaude 02af8463f6 Administer invitations in /admin/users (#1185)
* Add invited users with delete button to admin users page

Shows pending invitations per family below active users in /admin/users/.
Each invitation row has a red Delete button aligned with the role column.
Alt/option-clicking any Delete button changes all invitation button labels
to "Delete All" and destroys all pending invitations for that family.

- Add admin routes: DELETE /admin/invitations/:id and DELETE /admin/families/:id/invitations
- Add Admin::InvitationsController with destroy and destroy_all actions
- Load pending invitations grouped by family in users controller index
- Render invitation rows in a dashed-border tbody below active user rows
- Add admin-invitation-delete Stimulus controller for alt-click behavior
- Add i18n strings for invitation UI and flash messages

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Fix destroy_all using params[:id] from member route

The member route /admin/families/:id/invitations sets params[:id],
not params[:family_id], so Family.find was always receiving nil.

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Fix translation key in destroy_all to match locale

t(".success_all") looked up a nonexistent key; the locale defines
admin.invitations.destroy_all.success, so t(".success") is correct.

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Scope bulk delete to pending invitations and allow re-inviting emails

- destroy_all now uses family.invitations.pending.destroy_all so accepted
  and expired invitation history is preserved
- Replace blanket email uniqueness validation with a custom check scoped
  to pending invitations only, so the same email can be invited again
  after an invitation is deleted or expires

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Drop unconditional unique DB index on invitations(email, family_id)

The model-level uniqueness check was already scoped to pending
invitations, but the blanket unique index on (email, family_id)
still caused ActiveRecord::RecordNotUnique when re-inviting an
email that had any historical invitation record in the same family
(e.g. after an accepted invite or after an account deletion).

Replace it with no DB-level unique constraint — the
no_duplicate_pending_invitation_in_family model validation is the
sole enforcer and correctly scopes uniqueness to pending rows only.

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Replace blanket unique index with partial unique index on pending invitations

Instead of dropping the DB-level uniqueness constraint entirely, replace
the unconditional unique index on (email, family_id) with a partial unique
index scoped to WHERE accepted_at IS NULL. This enforces the invariant at
the DB layer (no two non-accepted invitations for the same email in a
family) while allowing re-invites once a prior invitation has been accepted.

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

* Fix migration version and make remove_index reversible

- Change Migration[8.0] to Migration[7.2] to match the rest of the codebase
- Pass column names to remove_index so Rails can reconstruct the old index on rollback

https://claude.ai/code/session_01F8WaH5TmtdUWwhHnVoQ6Gm

---------

Signed-off-by: Juan José Mata <juanjo.mata@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-03-14 11:32:33 +01:00
Juan José Mata 7fce804c89 Group users by family in /admin/users (#1139)
* Display user admins grouped

* Start family/groups collapsed

* Sort by number of transactions

* Display subscription status

* Fix tests

* Use Stimulus
2026-03-06 23:38:25 +01:00
Juan José Mata 06fedb34f3 Add new columns and sorting to admin users list (#1004)
* Add trial end date to admin users list

* Add new columns

* Regression
2026-02-16 20:10:14 +01:00
LPWandJosh Waldrep 320e087a22 Add support for displaying and managing legacy SSO providers (#628)
* feat: add support for displaying and managing legacy SSO providers

- Introduced UI section for environment/YAML-configured SSO providers.
- Added warnings and guidance on migrating legacy providers to database-backed configuration.
- Enhanced localization with new keys for legacy provider management.
- Updated form and toggle components for improved usability.

* Expand SSO documentation: add SAML 2.0 support, JIT provisioning settings, super-admin setup steps, audit logging, and user administration details.

* Update JIT provisioning docs: clarify role mapping behavior and add examples; note new `logout_idp` audit log event.

---------

Co-authored-by: Josh Waldrep <joshua.waldrep5+github@gmail.com>
2026-01-13 09:37:19 +01:00
Josh Waldrep 14993d871c feat: comprehensive SSO/OIDC upgrade with enterprise features
Multi-provider SSO support:
   - Database-backed SSO provider management with admin UI
   - Support for OpenID Connect, Google OAuth2, GitHub, and SAML 2.0
   - Flipper feature flag (db_sso_providers) for dynamic provider loading
   - ProviderLoader service for YAML or database configuration

   Admin functionality:
   - Admin::SsoProvidersController for CRUD operations
   - Admin::UsersController for super_admin role management
   - Pundit policies for authorization
   - Test connection endpoint for validating provider config

   User provisioning improvements:
   - JIT (just-in-time) account creation with configurable default role
   - Changed default JIT role from admin to member (security)
   - User attribute sync on each SSO login
   - Group/role mapping from IdP claims

   SSO identity management:
   - Settings::SsoIdentitiesController for users to manage connected accounts
   - Issuer validation for OIDC identities
   - Unlink protection when no password set

   Audit logging:
   - SsoAuditLog model tracking login, logout, link, unlink, JIT creation
   - Captures IP address, user agent, and metadata

   Advanced OIDC features:
   - Custom scopes per provider
   - Configurable prompt parameter (login, consent, select_account, none)
   - RP-initiated logout (federated logout to IdP)
   - id_token storage for logout

   SAML 2.0 support:
   - omniauth-saml gem integration
   - IdP metadata URL or manual configuration
   - Certificate and fingerprint validation
   - NameID format configuration
2026-01-03 17:56:42 -05:00