TPC-DS benchmarks

103 queries · six engines · one i4i.8xlarge · read from S3 · 100 GB and 1 TB · September 2026

Scale

Total query time

Sum of the best time per query, all 103 queries: best of two runs at 100 GB, one run at 1 TB. Lower is better.

Engine

100 GB · SF100

1 TB · SF1000

Databricks Standard

16.4

.

1,180 s

.

103.1 min
Spark

3.5.5

.

1,135 s

.

135.6 min
Comet

1.0.0

.

652 s

.

84.4 min
Gluten + Velox

1.7.0

.

557 s

.

46.9 min
Flarion

2026-09-12

.

386 s

.

38.3 min
Databricks Photon

16.4

.

362 s

.

30.4 min

Speedup

The selected baseline's total query time divided by each engine's, at the same scale. Above 1.0x is faster, below 1.0x is slower. Bars share one axis across both scales.

Baseline

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

baseline

.

1.00x

.

1.00x
Spark

3.5.5

.

1.04x

.

0.76x
Comet

1.0.0

.

1.81x

.

1.22x
Gluten + Velox

1.7.0

.

2.12x

.

2.20x
Flarion

2026-09-12

.

3.06x

.

2.69x
Databricks Photon

16.4

.

3.26x

.

3.39x

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

.

0.96x

.

1.32x
Spark

3.5.5

baseline

.

1.00x

.

1.00x
Comet

1.0.0

.

1.74x

.

1.61x
Gluten + Velox

1.7.0

.

2.04x

.

2.89x
Flarion

2026-09-12

.

2.94x

.

3.54x
Databricks Photon

16.4

.

3.14x

.

4.46x

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

.

0.55x

.

0.82x
Spark

3.5.5

.

0.57x

.

0.62x
Comet

1.0.0

baseline

.

1.00x

.

1.00x
Gluten + Velox

1.7.0

.

1.17x

.

1.80x
Flarion

2026-09-12

.

1.69x

.

2.20x
Databricks Photon

16.4

.

1.80x

.

2.78x

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

.

0.47x

.

0.46x
Spark

3.5.5

.

0.49x

.

0.35x
Comet

1.0.0

.

0.85x

.

0.56x
Gluten + Velox

1.7.0

baseline

.

1.00x

.

1.00x
Flarion

2026-09-12

.

1.44x

.

1.23x
Databricks Photon

16.4

.

1.54x

.

1.54x

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

.

0.33x

.

0.37x
Spark

3.5.5

.

0.34x

.

0.28x
Comet

1.0.0

.

0.59x

.

0.45x
Gluten + Velox

1.7.0

.

0.69x

.

0.82x
Flarion

2026-09-12

baseline

.

1.00x

.

1.00x
Databricks Photon

16.4

.

1.07x

.

1.26x

Engine

100 GB · SF100 · speedup

1 TB · SF1000 · speedup

Databricks Standard

16.4

.

0.31x

.

0.29x
Spark

3.5.5

.

0.32x

.

0.22x
Comet

1.0.0

.

0.55x

.

0.36x
Gluten + Velox

1.7.0

.

0.65x

.

0.65x
Flarion

2026-09-12

.

0.94x

.

0.79x
Databricks Photon

16.4

baseline

.

1.00x

.

1.00x

Total query time

Sum of best times, 103 queries. Lower is better.

Databricks Standard

16.4

.

1,180 s

?

Spark

3.5.5

.

1,135 s

?

Comet

1.0.0

.

652 s

?

Gluten + Velox

1.7.0

.

557 s

?

Flarion

2026-09-12

.

386 s

?

Databricks Photon

16.4

.

362 s

?

Spark

3.5.5

.

135.6 min

?

Databricks Standard

16.4

.

103.1 min

?

Comet

1.0.0

.

84.4 min

?

Gluten + Velox

1.7.0

.

46.9 min

?

Flarion

2026-09-12

.

38.3 min

?

Databricks Photon

16.4

.

30.4 min

?

Speedup

Databricks Standard's total query time divided by each engine's. Above 1.0x is faster, below 1.0x is slower.

Databricks Standard

16.4

baseline

.

1.00x

?

Spark

3.5.5

.

1.04x

?

Comet

1.0.0

.

1.81x

?

Gluten + Velox

1.7.0

.

2.12x

?

Flarion

2026-09-12

.

3.06x

?

Databricks Photon

16.4

.

3.26x

?

Spark

3.5.5

.

0.76x

?

Databricks Standard

16.4

baseline

.

1.00x

?

Comet

1.0.0

.

1.22x

?

Gluten + Velox

1.7.0

.

2.20x

?

Flarion

2026-09-12

.

2.69x

?

Databricks Photon

16.4

.

3.39x

?

Where an engine's own q72 plan is slower than Spark's, q72 is counted at Spark's time (Gluten and Comet at both scales). Every other engine keeps its own q72.

Benchmark figures, not customer results. Derived from TPC-DS under the TPC fair-use policy; not an audited TPC result.

What ran, and how it was configured

Shared environment

Machine

One i4i.8xlarge:

32 vCPU, 256 GB memory, 18.75 Gbps network, in us-west-2a.

The two Databricks rows ran on a single-node job cluster of the same instance type, in us-east-1.

Spark

Spark 3.5.5, local[32, 4].

One JVM per engine, Corretto 17 on EC2.

The Databricks rows run DBR 16.4 LTS (Spark 3.5.2), local[*, 4].

Data

The same parquet for every engine, read from S3 with no local cache.

100 GB scale, SF100, and 1 TB scale, SF1000.

Databricks reads the same parquet through a Unity Catalog volume, disk cache off.

Shared settings

64 shuffle partitions

AQE on

no table statistics or CBO

Equal memory budgets: Spark 120 GB heap (72 GB unified pool), every other engine 48 to 50 GB heap plus 72 GB off-heap

Engine configurations

Engine · build

Memory

Settings

Databricks Standard
DBR 16.4 LTS, Spark 3.5.2, Photon off

+

Apache Spark
3.5.5, vanilla, the correctness reference

+

Apache Comet
1.0.0 on Spark 3.5.5. Jar: comet-spark-spark3.5_2.12 (Spark 3.5, Scala 2.12) from Maven Central

+

Apache Gluten
1.7.0 on Spark 3.5.5, Velox backend. Jar: gluten-velox-bundle-spark3.5_2.12-linux_amd64 (Spark 3.5, Scala 2.12)

+

Flarion
Spark 3.5.5 · 2026-09-12

+

Databricks Photon
DBR 16.4 LTS, Spark 3.5.2, Photon on

+

Reproduce

Run it yourself

Every setting is on this page: the machine, the data, each engine's configuration and how the runs were timed. Spark, Comet and Gluten run from their public releases. The two Databricks rows need a Databricks workspace.

To see Flarion on your own Spark jobs, contact us.

Contact us

Method

Each engine ran the 103 queries

(TPC-DS v1.4, with q14, q23, q24 and q39 in both variants) in its own JVM against parquet on S3, with q3 first as an untimed warm-up.

Best of two runs at 100 GB, one run at 1 TB.

Each query had a cap of 600 s at 100 GB and 1,800 s at 1 TB. Totals sum the per-query times. Runs taken right after a cold start or a build were discarded, because the first query of a fresh JVM costs tens of seconds.

q72 is counted at Spark's time for Gluten and Comet.

It joins inventory to catalog_sales under a range condition. Gluten's maintainers recommend running it on vanilla Spark, and Comet's sort-merge join did not finish within the cap at 1 TB. Spark's measured time was 129.1 s at 100 GB and 81.5 s at 1 TB. Photon, Databricks Standard and Flarion beat Spark on q72 and keep their own numbers.

flarion icon

Run a free assessment

Now let's look at your
Spark workloads.

We start with an assessment of your Spark jobs, identify where Flarion could help, then validate the impact against your current Spark baseline.