The oceanbase driver install step reported success but never actually
installed oceanbase_py: --no-deps applied to `-e .[oceanbase]` blocks pip
from installing anything the extras marker pulls in, including
oceanbase_py itself, not just its conflicting transitive dependency.
Installing oceanbase_py as its own standalone package instead means
--no-deps only skips *its* dependencies, which is what was actually
intended. Confirmed via a manual workflow_dispatch run (nightly_only
dialects don't run on pull_request, so this needed a manual trigger to
catch at all).
Vertica dropped from this PR: the same workflow_dispatch run found
`vertica/vertica-ce` doesn't exist on Docker Hub. The only actively
maintained official image (`opentext/vertica-k8s`) is built to run under
the Vertica Kubernetes operator's orchestration, not as a standalone
single-container database -- a bare `docker run` likely won't bootstrap
a working instance on its own. Needs real investigation before it's
worth another attempt, same as Solr/IoTDB/TDengine/Parseable/Dremio
earlier in this series.
Stacked on feat/testcontainers-nightly-only-gating. All six extras already
existed in pyproject.toml. oceanbase and vertica run nightly_only: true
(heavy first-boot and a ~12GB RAM floor, respectively), so they don't run
per-PR; databend/risingwave/firebird/ydb run on every PR like the rest of
this suite.
- oceanbase_py pins sqlalchemy-utils<0.39, which conflicts outright with
Superset's own sqlalchemy-utils==0.42.1 pin -- kept out of the baseline
dev install (same reason as db2's ibm-db-sa) and installed on demand,
--no-deps, only for its own CI leg (it never actually imports
sqlalchemy_utils itself, so the version mismatch is inert at runtime).
- databend: connects to the local standalone image's builtin `root` user
(no password) with sslmode=disable, since Superset's default
encryption_parameters assume TLS the local image doesn't have.
- risingwave: RisingWave's storage engine checkpoints asynchronously --
a SELECT immediately after INSERT can see zero rows without an explicit
FLUSH (confirmed on a real instance). Uses the shared _pagination.py
helper's after_insert hook (originally added for CrateDB) to do that.
- firebird: sqlalchemy-firebird's driver is a pure-Python ctypes wrapper
(py3-none-any wheel, confirmed by downloading it directly) that
dynamically loads the native libfbclient from the host rather than
bundling it -- CI installs that system package on demand. Also confirms
in the test docstring that FirebirdEngineSpec's `limit_method =
LimitMethod.FETCH_MANY` (comment: "uses FIRST to limit") is stale
against the modern driver, which compiles real ROWS-based pagination.
- ydb: needed three real fixes to make a generic DockerContainer usable
at all. (1) YDB's gRPC client does endpoint discovery and reconnects to
whatever the server reports, which by default is the container's own
internal Docker hostname -- fixed by binding the same port on the host
as inside the container and advertising "localhost" as the container's
own hostname, so the discovered endpoint is actually reachable. (2) The
gRPC port opens before storage pools are fully initialized, so an early
CREATE TABLE fails; the fixture retries a real metadata.create_all()
probe rather than trusting the open port. (3) YDB rejects DDL inside an
explicit transaction ("Scheme operations cannot be executed inside
transaction") -- confirmed this only affects a raw text("CREATE
TABLE..."), not metadata.create_all()'s own DDL execution path, which
already does the right thing.
A job-level `if:` can't reference `matrix` at all -- only github/inputs/
needs/vars contexts are available there, confirmed by actionlint and by
this exact commit's own CI run failing outright with "This run likely
failed because of a workflow file issue" (zero jobs registered). Moves
the same condition onto the step that actually runs the tests instead,
where matrix access is already used successfully by the existing db2
install step.
A future dialect whose image is too heavy for per-PR CI (a multi-service
cluster, a many-GB image, a slow licensed installer) can set
nightly_only: true on its matrix entry to run only on the cron or a
manual workflow_dispatch, never on pull_request. No existing dialect
uses it yet -- this just lays the groundwork for candidates like SAP
HANA, Teradata, or Apache Druid.
Stacked on feat/testcontainers-more-dialects. All four extras already
existed in pyproject.toml, so this only wires up tests -- no new
optional-dependency groups needed.
- postgres/mysql: straight copies of the timescaledb/mariadb pattern
respectively, pointed at vanilla images instead of a fork/extension.
- clickhouse: connects over the container's HTTP port (8123), matching
clickhouse-connect (Superset's driver), not the native TCP port (9000)
the container's own docstring example uses. ClickHouse has no real
primary-key concept and clickhouse-connect's DDL compiler rejects
CREATE TABLE without an explicit engine, so _pagination.py gained an
optional extra_table_args hook to pass MergeTree(order_by=...). Also
works around a pre-existing quirk in db_engine_specs/clickhouse.py:
its module-level type-formatting setup dereferences current_app.config,
so importing it outside a Flask app context raises RuntimeError --
tests/unit_tests/db_engine_specs/test_clickhouse.py already works
around this with per-test local imports, but that suite also benefits
from an autouse app_context fixture this suite doesn't have, so this
test pushes one explicitly around the one-time import.
- starrocks: no dedicated testcontainers module, so a generic
DockerContainer against the official allin1-ubuntu image (FE+BE in one
container). Not verified locally (multiple-GB image, skipped to keep
local Docker load low per session guidance); the fixture retries its
first connection since the query port can accept TCP before StarRocks'
query engine is fully initialized.
- drop `db2` from the `testcontainers[...]` extra in development.in/.txt:
it transitively pulls `ibm-db-sa`/`ibm-db` into the baseline dev lockfile,
which has no Linux arm64 wheel and breaks the multi-platform dev Docker
image build; the db2 CI leg already installs it on demand separately.
- mariadb: preserve a remote Docker daemon's real host, only rewrite the
literal "localhost" case to 127.0.0.1.
- testcontainers.yml: widen pull_request paths to superset/db_engine_specs/**,
pyproject.toml, and the requirements manifests so a driver/lockfile-only
change can't bypass this coverage.
- mongodb: use Core `select(...).limit().offset()` instead of a literal SQL
string, so the pagination test actually exercises the dialect's own
compilation.
- monetdb/mongodb/yugabytedb: assert generic_type/sqla_type on the mapped
column spec, not just that a spec was returned.
- monetdb: wait for the exposed port (not just a log line that predates
`monetdbd start -n`) before considering the container ready.
- timescaledb: pin the image to a specific tag instead of the mutable
`latest-pg16`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds five more dialects to the testcontainers suite, stacked on top of the
7-dialect pilot in feat/testcontainers-nightly-pilot: mariadb, timescaledb,
yugabytedb, monetdb, mongodb.
mariadb, timescaledb, and yugabytedb are all wire-compatible with an
existing base dialect (MySQL and Postgres respectively), so they reuse
testcontainers' MySqlContainer/PostgresContainer classes pointed at a
different image rather than needing new container-class wiring.
yugabytedb specifically cannot reuse PostgresContainer's built-in
readiness check, though: that execs `psql`, which the yugabyte image
doesn't ship (only its own `ysqlsh`) -- uses a generic DockerContainer
instead, started via `yugabyted start` and waiting on its own final
startup log line.
monetdb has no native testcontainers module; uses a generic DockerContainer
with the documented MDB_* environment variables. Publishes an amd64-only
image (confirmed running under Rosetta/QEMU emulation on Apple Silicon,
unlike CrateDB's harder x86-64-v3 CPU requirement).
mongodb needed a different data-setup approach, like elasticsearch before
it: documents get inserted via the native pymongo driver, not SQL INSERT,
since MongoDB is schemaless and Superset talks to it through pymongosql
(a SQL-to-MongoDB translation layer requiring a `?mode=superset` query
param). Two real bugs surfaced writing this one: testcontainers'
MongoDbContainer.get_connection_url() has no database path or query
string at all, so naively appending "&mode=superset" glued directly onto
the port number instead of starting a query string; and the root user
MongoDbContainer creates lives in the `admin` database, so connecting
with a different default database in the URL requires authSource=admin
or authentication fails outright. Confirmed pymongosql supports OFFSET
(maps to MongoDB's native `skip`), unlike Elasticsearch's SQL layer.
mariadb could not be verified locally in this environment: mysqlclient
(MySQLdb) has a pre-existing, unrelated native-library linking issue
against this machine's Homebrew-installed libmysqlclient. CI installs it
via apt on Linux, where this does not occur -- same accepted pattern
already used for crate/mssql/db2 in the base branch.
- requirements/development.in + testcontainers.yml: scope the `db2` extra
(`ibm-db-sa`/`ibm-db`) out of the baseline dev install. `ibm-db` ships no
Linux arm64 wheel, so bundling it there broke the multi-platform dev
Docker image build on push. The testcontainers CI job now installs it
directly, only for its own db2 matrix leg.
- cockroachdb.py: list `psycopg2-binary` alongside `sqlalchemy-cockroachdb`
in the `pypi_packages` metadata, since a plain `cockroachdb://` URL can't
connect without a DBAPI and sqlalchemy-cockroachdb doesn't install one.
- test_cockroachdb.py/test_crate.py/test_trino.py: fix stale docstring
references to the old `nightly-testcontainers.yml` filename.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- pyproject.toml: pin psycopg2-binary alongside sqlalchemy-cockroachdb --
the latter declares no DBAPI dependency of its own, so the documented
`apache-superset[cockroachdb]` install couldn't actually connect.
- testcontainers.yml: scope the concurrency group by ref so a PR run and
the nightly cron (or two different PRs) no longer cancel each other.
- test_cockroachdb.py: assert the actual generic/SQLAlchemy type, matching
the Trino test, instead of only checking a column spec was found.
- pytest.ini + new `testcontainers` marker + _driver.py: exclude
tests/testcontainers/ from a plain `pytest` run by default (it needs
Docker), while the dedicated CI job now sets
SUPERSET_TESTCONTAINERS_STRICT so a broken/missing driver import fails
that job instead of silently skipping to a green, zero-tests-run result.
- UPDATING.md: the migration note now says to uninstall the old
`cockroachdb` package outright, since reinstalling the extra alone can
leave both packages registering the same dialect entry point.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Adds four more dialects to the testcontainers suite (cockroachdb, crate,
trino from the initial pilot): mssql, oracle, db2, elasticsearch. All four
have native testcontainers-python container classes.
Restructures the workflow from one job running the whole suite to a
matrix, one job per dialect, running in parallel. A single slow container
would otherwise inflate wall-clock time for every dialect, not just its
own -- matrixing bounds total suite time by the slowest dialect instead of
the sum of all of them. Also renames the workflow file/name from
"Nightly-Testcontainers" to "Testcontainers" now that it runs on
pull_request (scoped via `paths`) in addition to the nightly cron.
Elasticsearch needed a different data-setup approach than the SQL-native
dialects: indices/documents get created via its REST API, not SQL INSERT,
matching how Superset actually encounters Elasticsearch in practice.
Confirmed empirically that Elasticsearch's SQL layer has no OFFSET support
at all (a real protocol limitation, already correctly documented via
ElasticSearchEngineSpec.supports_offset = False) and adjusted that
dialect's pagination test accordingly -- LIMIT/ORDER BY only, no OFFSET.
mssql and db2 could not be verified locally (no arm64 images for either;
this environment is Apple Silicon), same situation as crate's amd64-only
image from the initial pilot. Both are written against verified library
source (dialect names, connection URL construction) and will get their
first real execution on CI.
Adds a scoped pull_request trigger so this workflow runs on this PR
itself, to get real GHA runner timing before deciding whether/how to
adopt this pattern more broadly. Also adds the actions-timeline job
(same pattern as superset-python-presto-hive.yml) to visualize per-step
duration.
Revert the pull_request trigger before merge -- it's here to measure
cost, not to become a permanent merge-blocking check.
Pilot for real-container testing of db_engine_specs against actual
databases (CockroachDB, CrateDB, Trino), via testcontainers-python.
Existing unit tests mock the driver/dialect layer entirely, which cannot
catch real SQL-compilation or type-mapping bugs -- e.g. #42899, where
Trino emitted OFFSET before LIMIT for paginated queries.
Runs nightly (.github/workflows/nightly-testcontainers.yml), not on every
merge: container pulls and startup are slower and more flake-prone than
the existing mocked unit tests, and this measures that cost/signal
tradeoff before considering wider adoption.
Building this surfaced two real bugs, fixed/documented separately:
- The cockroachdb dialect was completely broken under SQLAlchemy 2.0 due
to a dead upstream package -- fixed in #43501.
- testcontainers-python's TrinoContainer.get_connection_url() returns the
container-internal port instead of the Docker-mapped host port; worked
around locally, filed upstream.