SkaleData Docs
Superset Image

Overview

The SkaleData Superset image: preinstalled database drivers, how to extend it, and how to deploy a custom image.

SkaleData publishes a maintained Superset image at ghcr.io/skaledata/superset. It's the official apache/superset image with drivers for common warehouses and databases, so you can connect to the databases in Preinstalled databases without building an image of your own.

New SkaleData Superset instances run this image by default. An existing instance moves to it the next time it's applied or restarted: a Save & Apply, a Restart, or a cluster apply.

This page describes ghcr.io/skaledata/superset:6.1.0-dc3cdfa, the default image.

What's in the image

  • apache/superset 6.1.0, unchanged. Adding the drivers doesn't change the version of any package that the upstream image already has.
  • The database drivers in Preinstalled databases.
  • The shared libraries those drivers load: the MariaDB client for MySQL, Cyrus SASL for Hive and Spark SQL, and the Firebird client library.
  • The packages that Superset's built-in Model Context Protocol (MCP) server needs, so that the superset mcp run command works: fastmcp 3.2.4, mcp 1.29.1, and their dependencies, such as starlette 1.3.1 and uvicorn. They aren't apache-superset extras, so the tables on this page don't list them. The image doesn't start the MCP server.

The image is published for linux/amd64 only. Cluster nodes are amd64. On a computer with Apple silicon, Docker runs the image under emulation.

Tags

TagMoves?Use it for
6.1.0Yes. It moves to each new release of the image for Superset 6.1.0.Custom images that should pick up driver fixes when you rebuild.
6.1.0-<sha7>No. <sha7> is the commit the image was built from.Reproducible builds.
latestYes. It follows the newest supported Superset version.Trying the image out. Don't build production images on it.

The default image is 6.1.0-dc3cdfa, pinned by digest:

ghcr.io/skaledata/superset:6.1.0-dc3cdfa@sha256:b2328a3d6c564773414993df8bcf27cb696d22fa03559cd12c9c2ea2b216db34

The package is public, so pulling it doesn't need credentials:

docker pull --platform linux/amd64 ghcr.io/skaledata/superset:6.1.0-dc3cdfa

Preinstalled databases

Superset lists these databases in its Connect a database dialog. The extra column names the apache-superset extra that the image installs for each one. The driver packages column names the Python packages that provide the connection.

DatabaseExtraDriver packages
Amazon Athenaathenapyathena
Amazon DynamoDBdynamodbpydynamodb
Amazon Redshiftredshiftsqlalchemy-redshift, psycopg2-binary
Apache Dorisdorispydoris, mysqlclient
Apache Drilldrillsqlalchemy-drill
Apache Druiddruidpydruid
Apache Hivehivepyhive, thrift, thrift-sasl
Apache Impalaimpalaimpyla
Apache Kylinkylinkylinpy
Apache Pinotpinotpinotdb
Apache Spark SQLsparkpyhive, thrift
Aurora MySQLmysqlmysqlclient
Aurora MySQL (Data API)aurora-data-apipreset-sqlalchemy-aurora-data-api
Aurora PostgreSQLpostgrespsycopg2-binary
Aurora PostgreSQL (Data API)aurora-data-apipreset-sqlalchemy-aurora-data-api
ClickHouseclickhouseclickhouse-connect
CockroachDBcockroachdbcockroachdb, psycopg2-binary
CrateDBcratesqlalchemy-cratedb
Databricksdatabricksdatabricks-sqlalchemy, databricks-sql-connector
Databricks (legacy)databricksdatabricks-sqlalchemy, databricks-sql-connector
Databricks Interactive Clusterdatabricksdatabricks-sqlalchemy, databricks-sql-connector
Databricks SQL Endpointdatabricksdatabricks-sqlalchemy, databricks-sql-connector
Dremiodremiosqlalchemy-dremio
DuckDBmotherduckduckdb-engine, duckdb
Elasticsearchelasticsearchelasticsearch-dbapi
Firebirdfirebirdsqlalchemy-firebird, fdb
Fireboltfireboltfirebolt-sqlalchemy
Google BigQuerybigquerysqlalchemy-bigquery, google-cloud-bigquery, pandas-gbq
Google Sheetsgsheetsshillelagh
MotherDuckmotherduckduckdb-engine, duckdb
MySQLmysqlmysqlclient
OpenSearch (OpenDistro)elasticsearchelasticsearch-dbapi
PostgreSQLpostgrespsycopg2-binary
Prestoprestopyhive
SQLiteIncluded with SupersetPython's sqlite3 module
ShillelaghIncluded with Supersetshillelagh
SingleStoresinglestoresqlalchemy-singlestoredb, singlestoredb
Snowflakesnowflakesnowflake-sqlalchemy
StarRocksstarrocksstarrocks, pymysql
Trinotrinotrino
Verticaverticasqlalchemy-vertica-python, vertica-python

The image also installs this driver, which Superset doesn't list in the dialog:

DatabaseExtraDriver packages
Microsoft SQL Servermssqlpymssql

Some notes on the list:

  • SQLite and Shillelagh appear in the list, but Superset blocks connections to them by default. Its PREVENT_UNSAFE_DB_CONNECTIONS setting is on, because both can read files on the Superset server.
  • DuckDB comes with the motherduck extra, which installs Superset's duckdb extra.
  • Databricks: Superset lists four Databricks entries because they share one SQLAlchemy backend. Use Databricks, which connects with a databricks:// URI. The other three default to databricks+connector://, databricks+pyhive://, and databricks+pyodbc:// URIs, and the image doesn't include those drivers.

Connect to Microsoft SQL Server

The image installs pymssql, but Superset lists Microsoft SQL Server in the Connect a database dialog only when pyodbc is installed. pyodbc needs the Microsoft ODBC driver, which the image doesn't include. To connect through pymssql, enter a SQLAlchemy URI:

  1. In Superset, click + > Data > Connect database. You can also go to Settings > Database Connections and click + Database.

  2. Under Or choose from a list of other databases we support, open the Supported databases list and select Other.

  3. In Display Name, enter a name for the connection.

  4. In SQLAlchemy URI, enter a mssql+pymssql:// URI:

    mssql+pymssql://<username>:<password>@<host>:<port>/<database>
  5. Click Test connection, and then click Connect.

For the URI's options, see the SQLAlchemy SQL Server dialect documentation.

Databases that aren't included

The image doesn't include drivers for these databases:

DatabaseSuperset extraWhy it's left out
IBM Db2db2The ibm-db driver bundles IBM's CLI driver. Its license, the IBM International Program License Agreement, restricts redistribution and hosting.
SAP HANAhanaThe SAP Developer License Agreement bars making the hdbcli driver available to third parties. You can add it yourself.
TeradatateradataThe Teradata License Agreement bars distributing or embedding the teradatasql driver without Teradata's written consent. You can add it yourself.
Oracleoraclecx-Oracle can't connect without Oracle Instant Client, a licensed system library.
Azure SynapseNoneIt needs Microsoft's ODBC driver, a licensed system library.
Databricks over ODBCNonedatabricks+pyodbc:// needs the Databricks (Simba Spark) ODBC driver, a licensed system library. Use the Databricks entry instead.
ExasolexasolSuperset's exasol extra requires sqlalchemy-exasol 2.x, and no 2.x release supports SQLAlchemy 1.4, which Superset 6.1 runs on.
Airtable, MongoDBNoneSuperset has no extra for them.
RocksetNoneRockset shut down its service in 2024.

Add the SAP HANA or Teradata driver

You can install either driver from PyPI through your own requirements.txt, as described in Extend the image. When you do, you download the driver yourself, and you accept the vendor's license directly: the SAP Developer License Agreement for hdbcli, or the Teradata License Agreement for teradatasql and teradatasqlalchemy. Read the license, and decide whether it allows your use, before you build.

For SAP HANA, add the hdbcli driver and the sqlalchemy-hana dialect:

# requirements.txt
hdbcli
sqlalchemy-hana==0.4.0

For Teradata, add the teradatasql driver and the teradatasqlalchemy dialect:

# requirements.txt
teradatasql
teradatasqlalchemy==20.0.0.2

Superset 6.1 runs on SQLAlchemy 1.4, which is why both dialects are pinned:

  • sqlalchemy-hana 0.4.0 is the version that Superset's own hana extra pins. Without the pin, the install picks a newer release that upgrades SQLAlchemy to 2.0, and Superset stops working.
  • teradatasqlalchemy releases after 20.0.0.2 need SQLAlchemy 2.0. On SQLAlchemy 1.4, their dialect fails to load, and Superset leaves Teradata out of its database list without showing an error.

After the build, SAP HANA or Teradata appears in the Connect a database dialog.

Extend the image

To add Python packages or Debian packages, build your own image on top of this one. Put a Dockerfile in a directory, and add a requirements.txt file, a packages.txt file, or both, next to it:

.
├── Dockerfile
├── packages.txt
└── requirements.txt

The Dockerfile needs one line:

FROM ghcr.io/skaledata/superset:6.1.0

packages.txt lists Debian packages, one per line:

# Debian packages, one per line. apt-get installs them as root.
postgresql-client  # psql, for testing connections from a pod

requirements.txt lists Python packages, in the usual requirements file format:

# Python packages, one requirement per line.
sqlalchemy-kusto==3.1.1  # Azure Data Explorer (Kusto)

Build it for linux/amd64:

docker build --platform linux/amd64 -t my-superset:6.1.0-build1 .

The base image declares ONBUILD triggers. They run at the start of your build, right after your FROM line and before your own instructions:

  1. If packages.txt exists, apt-get installs each package in it, as root.
  2. If requirements.txt exists, uv installs it into Superset's virtual environment, /app/.venv, as root.
  3. The build switches back to the superset user. The image runs as superset.

Both files are optional. Blank lines and # comments, whether they fill a whole line or follow a package name, are ignored. A file that contains only comments installs nothing.

The image has no compiler. If a package you add builds from source, add build-essential, and any -dev headers the package needs, to packages.txt.

requirements.txt installs without constraints

The ONBUILD step installs requirements.txt without version constraints. A package you add can upgrade or downgrade a package that Superset depends on, such as SQLAlchemy, Flask-AppBuilder, or even apache-superset itself. After a change like that, Superset can fail to start. Check the build output for changed or removed packages before you deploy.

Keep Superset's dependencies at their versions

To make the install fail instead of changing a package that the image already has, install your packages with the image's package list as a constraint. /app/skaledata/image-freeze.txt lists every package in the image and its version, except apache-superset and superset-core. Those two install from /app, not from PyPI, so the constraint doesn't pin Superset itself. Check the build output for a change to apache-superset.

FROM ghcr.io/skaledata/superset:6.1.0

USER root
ENV HOME=/root
COPY my-requirements.txt /tmp/my-requirements.txt
RUN uv pip install --python /app/.venv/bin/python --no-cache \
      --constraint /app/skaledata/image-freeze.txt \
      --requirement /tmp/my-requirements.txt
ENV HOME=/app/superset_home
USER superset
  • Name the file anything but requirements.txt, such as my-requirements.txt. The ONBUILD step installs a file named requirements.txt without constraints before your RUN step runs, so a constraint in your RUN step comes too late for it.
  • The install runs as root, because Superset's virtual environment belongs to root. USER superset switches back afterward.
  • The HOME lines keep files that root creates during the build out of the superset user's home directory.

If a package needs a different version of a package in the image, the build fails with an error like this one:

× No solution found when resolving dependencies:
╰─▶ Because you require humanize==4.13.0 and humanize==4.12.3, we can
    conclude that your requirements are unsatisfiable.

Deploy a custom image

The deploy script, deploy.sh, builds your image, pushes it to your cluster's container registry, and deploys it to one Superset app. You don't need cloud credentials; the script gets short-lived registry credentials from the SkaleData API.

Before you begin

  • Docker, jq, and curl 7.76 or later.
  • A SkaleData API key with the apps:deploy scope, or the full scope. See API key scopes. If you call these endpoints with a console session instead of a key, you need the Operator role or higher; see Roles and permissions.
  • The cluster's ID. It's the last part of the cluster's page URL in the console, /dashboard/clusters/<cluster-id>.
  • The Superset app's name. You can leave it out if the cluster has only one Superset app.

Get the script and run it

  1. Download the script for your Superset app:

    export SKALEDATA_API_KEY=sdk_...
    curl -sS --fail-with-body \
      -H "Authorization: Bearer $SKALEDATA_API_KEY" \
      -o deploy.sh \
      "https://api.skaledata.com/clusters/<cluster-id>/deploy-script?app_type=superset&app_name=<app-name>"
    chmod +x deploy.sh
  2. Put deploy.sh in the directory with your Dockerfile. The script's header has an example Dockerfile that starts FROM the exact image tag your cluster runs.

  3. Run the script from that directory:

    ./deploy.sh

The script does the following:

  1. Gets registry credentials from POST /clusters/<cluster-id>/registry-token and signs in to the registry with docker login.
  2. Builds the image for linux/amd64, even on a computer with Apple silicon.
  3. Pushes it as <registry>/<app-name>:<tag>. By default, the tag is the Superset version and a UTC timestamp, such as 6.1.0-20260928171233, so each run pushes a new tag. To choose the tag, set IMAGE_TAG, for example IMAGE_TAG=6.1.0-build42 ./deploy.sh.
  4. Reads the digest the registry reports for the push. If it can't read a sha256: digest, it stops before it deploys anything.
  5. Calls POST /clusters/<cluster-id>/deploy-image with app_type, app_name, image_tag, and image_digest.

The API responds like this:

{
  "status": "deploying",
  "image": "<registry>/<app-name>:6.1.0-20260928171233@sha256:<digest>",
  "app": "<app-name>"
}

The deploy pins the pushed digest. If you push the same tag again with new content and deploy it, the pods still move to the new content.

Tag rules

The API checks the tag before it deploys, and returns 400 Bad Request when a rule fails:

  • No latest or current. Each deploy's tag names the build it runs, so you can trace what an app runs back to a build. Use a unique tag for each deploy, such as a git SHA or a timestamp.

  • The Superset version must match. The API reads a tag that starts with two dot-separated numbers, with or without a leading v, as a Superset version. For example, 6.1.0-build42, v6.1.0, and 6.1 all read as Superset 6.1. That version must match your cluster's Superset major and minor version, 6.1, because an image on another Superset version migrates the metadata database. A date such as 2026.09.28 also reads as a version, 2026.9, so it returns 400. For example, 6.2.0-build1 fails with:

    The tag '6.2.0-build1' looks like Superset 6.2, but this cluster runs Superset 6.1. A different version migrates the metadata database, so build FROM the base image for Superset 6.1.
  • A tag without a version is allowed, with a warning. A tag such as a git SHA or 20260928-build1 deploys, and the response includes a warning field that reminds you to build FROM the base image for your cluster's Superset version.

  • The tag must be a valid Docker tag: up to 128 letters, digits, underscores, periods, and hyphens, not starting with a period or hyphen.

  • image_digest is required for Superset, as sha256: followed by 64 lowercase hexadecimal digits. deploy.sh sends it for you.

What happens after a deploy

Every deploy runs a Helm upgrade of the Superset app. It moves the web server, the Celery worker, Celery beat, and the database-migration job to the new image. The upgrade takes several minutes. In a test on 2026-09-28, the Helm job started about 4 to 6 minutes after the request, and then the pods rolled in about 1.5 minutes.

deploy-image returns 409 Conflict, and changes nothing, in these cases:

  • The cluster isn't in the ready state.
  • The cluster's container registry isn't provisioned.
  • Superset isn't enabled on the cluster.
  • The cluster has more than one Superset app, and the request has no app_name.
  • The app is provisioning, restarting, or being deleted. A deploy sets the app to provisioning until its Helm job finishes, so running deploy.sh again before the previous deploy finishes returns 409. The script pushes the image before it calls the API, so that push lands in the registry but isn't deployed.

Go back to the default image

To drop your custom image and run the SkaleData default image again, send reset_to_default. The curl command is also in the header of deploy.sh:

curl -sS --fail-with-body -X POST \
  "https://api.skaledata.com/clusters/<cluster-id>/deploy-image" \
  -H "Authorization: Bearer $SKALEDATA_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"app_type": "superset", "app_name": "<app-name>", "reset_to_default": true}'

Don't send image_tag or image_digest with reset_to_default. The reset also runs a Helm upgrade, and the response's image is the default image.

The console doesn't show a Superset app's custom image. To check which image an app runs, call GET /applications/<app-id> with a key that has the apps:read or full scope. <app-id> is the last part of the app's page URL in the console, /dashboard/applications/<app-id>. In the response:

  • config.last_deployed_image is the image that the last deploy or reset recorded, pinned by digest.
  • config.superset_custom_image holds the custom image. After a reset, that field is gone, and the app runs the default image.

On this page