TPC-DS benchmarks
103 queries · six engines · one i4i.8xlarge · read from S3 · 100 GB and 1 TB · September 2026
Total query time
Sum of the best time per query, all 103 queries: best of two runs at 100 GB, one run at 1 TB. Lower is better.
Engine
100 GB · SF100
1 TB · SF1000
Speedup
The selected baseline's total query time divided by each engine's, at the same scale. Above 1.0x is faster, below 1.0x is slower. Bars share one axis across both scales.
Engine
100 GB · SF100 · speedup
1 TB · SF1000 · speedup
Engine
100 GB · SF100 · speedup
1 TB · SF1000 · speedup
Engine
100 GB · SF100 · speedup
1 TB · SF1000 · speedup
Engine
100 GB · SF100 · speedup
1 TB · SF1000 · speedup
Engine
100 GB · SF100 · speedup
1 TB · SF1000 · speedup
Total query time
Sum of best times, 103 queries. Lower is better.
Speedup
Databricks Standard's total query time divided by each engine's. Above 1.0x is faster, below 1.0x is slower.
Where an engine's own q72 plan is slower than Spark's, q72 is counted at Spark's time (Gluten and Comet at both scales). Every other engine keeps its own q72.
Benchmark figures, not customer results. Derived from TPC-DS under the TPC fair-use policy; not an audited TPC result.
Setup
What ran, and how it was configured
Shared environment
Machine
One i4i.8xlarge:
32 vCPU, 256 GB memory, 18.75 Gbps network, in us-west-2a.
The two Databricks rows ran on a single-node job cluster of the same instance type, in us-east-1.
Spark
Spark 3.5.5, local[32, 4].
One JVM per engine, Corretto 17 on EC2.
The Databricks rows run DBR 16.4 LTS (Spark 3.5.2), local[*, 4].
Data
The same parquet for every engine, read from S3 with no local cache.
100 GB scale, SF100, and 1 TB scale, SF1000.
Databricks reads the same parquet through a Unity Catalog volume, disk cache off.
Shared settings
64 shuffle partitions
AQE on
no table statistics or CBO
Equal memory budgets: Spark 120 GB heap (72 GB unified pool), every other engine 48 to 50 GB heap plus 72 GB off-heap
Engine configurations
Engine · build
Memory
Settings
+
+
+
+
+
+
Reproduce
Run it yourself
Every setting is on this page: the machine, the data, each engine's configuration and how the runs were timed. Spark, Comet and Gluten run from their public releases. The two Databricks rows need a Databricks workspace.
To see Flarion on your own Spark jobs, contact us.
Contact usMethod
Each engine ran the 103 queries
(TPC-DS v1.4, with q14, q23, q24 and q39 in both variants) in its own JVM against parquet on S3, with q3 first as an untimed warm-up.
Best of two runs at 100 GB, one run at 1 TB.
Each query had a cap of 600 s at 100 GB and 1,800 s at 1 TB. Totals sum the per-query times. Runs taken right after a cold start or a build were discarded, because the first query of a fresh JVM costs tens of seconds.
q72 is counted at Spark's time for Gluten and Comet.
It joins inventory to catalog_sales under a range condition. Gluten's maintainers recommend running it on vanilla Spark, and Comet's sort-merge join did not finish within the cap at 1 TB. Spark's measured time was 129.1 s at 100 GB and 81.5 s at 1 TB. Photon, Databricks Standard and Flarion beat Spark on q72 and keep their own numbers.
Run a free assessment
Now let's look at your
Spark workloads.
We start with an assessment of your Spark jobs, identify where Flarion could help, then validate the impact against your current Spark baseline.