A Drop-In Execution Engine for Your Analytics Workloads

The same workloads finish faster, cost less, use less memory, and fail less. No code changes, nothing new for your team to operate.
Integrates with your analytics platform. No lock-in.

Same Job, Different Engine

Flarion drops into your existing stack, upgrades the query plan, and runs it on a native engine instead of the JVM.
Your cluster: Spark plans the job as usual
Submitted plan: 8 steps · 2 shuffles
ScanParquet #1
→
Exchange #1
→
ScanParquet #2
→
Exchange #2
→
SortMergeJoin
→
BatchEvalPython
→
HashAggregate
→
WriteParquet
flarion icon
Flarion
upgrades the plan
The upgraded plan runs natively
Flarion’s plan: 6 steps · 1 shuffle
ScanParquet
→
Exchange
→
SortMergeJoin
→
BatchEvalPython
→
HashAggregate
→
WriteParquet
→
→
2 fewer steps · 1 fewer shuffle
Flarion Accelerated
UDF Runs Unchanged
Read
Write
Your data lake · unchanged
Table formats
File formats
Parquet
Avro
JSON

Built on Apache Arrow and DataFusion

Flarion's optimizer, execution engine, and memory management are built on proven technology: Apache Arrow and DataFusion, implemented in Rust.
Apache Arrow

Columnar memory format

Each column sits contiguous in memory, so the query reads exactly what it asked for. Every modern analytical engine is columnar for this reason.

DataFusion

Query engine framework

An Apache top-level query engine with vectorized execution: it runs queries in parallel across cores and keeps data columnar the whole way.

Rust

Native, compiled operators

Compiled native operators, no garbage collector. The engine manages its own memory and frees it the moment an operation finishes.

Hundreds of Apache Commits

The team behind Flarion

Our team contributes to Apache Arrow and DataFusion, with hundreds of merged commits across Apache projects.

Where the Speed Comes From

An improved query plan, columnar data, native execution, and workload caching. Four levers, with the metrics to see each at work.
Fewer steps, so the job finishes sooner.
Filtering and sorting run faster on columns.
More work per cycle, no garbage collector.
Work you’ve run before doesn’t run again.
Fewer steps · One less shuffle
Before the job runs, Flarion upgrades the plan, merging steps and cutting a shuffle between stages. Same job, less to execute.
Submitted plan: 7 steps · 2 shuffles
ScanParquet #1
→
Filter
→
Exchange #1
→
ScanParquet #2
→
HashAggregate
→
Exchange #2
→
WriteParquet
Upgraded Flarion plan: 5 steps · 1 shuffle
ScanParquet
→
Filter
→
Exchange
→
HashAggregate
→
WriteParquet
→
Skipped Step
→
Skipped Step
Columnar data · Filter and sort without walking rows
A row layout packs all of a record's fields together on disk, so a query that needs one column has to read through all of them and discard most of what it touched. A column layout stores columns contiguously: the query reads the values it asked for and nothing else. That is why every modern analytical engine is columnar.
The Job: Get every zip code
Row layout · Reads 16 fields to keep 4
Row layout · Reads 16 fields to keep 4
Column layout · Reads only the zip column
Column layout · Reads only the zip column
Native · Lower CPU, flat memory
Flarion's engine is native code running close to the hardware. It uses resources sparingly and releases them as soon as the work is done. Lower CPU, flat memory, no GC pauses. The same job done sooner.
Spark UI · Details for Stage 4 (Attempt 0)
Spark UI · Details for Stage 4 (Attempt 0)
Details for Stage 4 · With Flarion
Details for Stage 4 · With Flarion
Caching · Only new work executes
Flarion caches data and results it has already computed. A later run that needs the same data is served from cache instead of recomputing it. Only genuinely new work executes.
Run 1 · Everything computes
ScanParquet
→
SortMergeJoin
→
HashAggregate
→
ScanParquet
•CASHED•
→
WriteParquet
Run 2 · The same job, run again
ScanParquet
•CASHED•
→
SortMergeJoin
•CASHED•
→
HashAggregate
•CASHED•
→
WriteParquet
•NEW WORK•

Speed Shows Up on the Bill

Fewer minutes on ephemeral clusters, fewer nodes on long-running ones.
3× Faster
60% Lower Cost

Spin up per job, tear down

The cluster ends when the job does.

Finish sooner and the cluster shuts down sooner. You pay for fewer minutes, automatically.

Diagram comparing two horizontal bars labeled 'Today' and 'With Flarion' showing that with Flarion, the bar is shorter by some minutes, indicating you stop paying sooner on autoscaled clusters.

Always-on, shared

Same jobs, a smaller cluster.

The jobs finish with headroom. Telemetry shows how many nodes you can remove.

Comparison showing eight empty squares labeled Today versus five filled yellow squares labeled With Flarion, indicating same work is done with fewer nodes.

Jobs Fail Less

Flarion is built to finish jobs. It manages memory natively, recognizes cloud and on-premises errors, and recovers on its own where it can.
Scatter plot comparing run durations under the same workload with and without Flarion enabled. On the left side, 71 Spark runs are shown as light blue dots with durations mostly between 30 minutes to 2 hours and some failed runs marked with orange crossed circles. The average duration for Spark runs is about 1 hour 3 minutes. On the right side, 64 Flarion runs are shown as yellow dots clustered around durations of about 30 minutes to 1 hour. The average duration for Flarion runs is about 37 minutes 2 seconds, indicating faster completion times and no failures.
Robust Memory Handling

Memory is freed the moment an operation finishes, so memory-bound jobs that fail on Spark can complete.

Automatic Recovery

Flarion recognizes known failures as they happen and, where it can, recovers and keeps the job running.

Evidence-Backed Fixes

When a job does fail, you get the evidence and a concrete fix, not a log to decode.

See Every Run

Flarion’s telemetry gives you all the metrics you need to report on Spark performance. Jump straight to any run: runtimes, errors, and detail down to the operator level.
Dashboard showing Flarion runs for the last 7 days with workloads, status, duration, native coverage, and Spark percentage; detailed view of nightly-compaction run highlighting write parquet stage, 11 GB peak memory, no stragglers, all stages ok, and telemetry stage timeline with read clicks, filter region, and project columns metrics.

Deploy Anywhere, in Minutes

Flarion deploys as one JAR plus config on the platform you already run. Setup in 6 minutes: one config change, no new permissions, no new infrastructure.
# add to your Spark
configuration

spark.jars
flarion-data-engine.jar
spark.sql.extensions
flarion.extensions.DataEngine
One JAR plus two config lines.
Databricks

Deploy via Init Scripts for runtime optimization.

Amazon EMR

Deployed as a bootstrap action.

Google Cloud Dataproc

Configured with initialization actions.

Azure HDInsight

Integrated via script actions for enhanced performance.

Spark on Kubernetes

Deploy with Helm charts or Spark operator modifications.

On-Premises

Install on Spark nodes using tools like Ansible or Chef, optimizing SQL operations.

Your Data Stays Within Your Infrastructure

Nothing leaves your environment except optional telemetry.
Data Privacy

No agent or service installed.

  • No access to your business data
  • Optional telemetry, anonymized, and encrypted
Enterprise Security

Secure by design.

  • SOC 2 Type II certified
  • Regular security audits
Deployment Control

No access to your environment.

  • One JAR, deployed by your team under your existing role
  • Air-gapped setups supported

Built for Your Stack

Flarion covers key frameworks, data lakes, file formats, and workloads.

Scala Spark · PySpark · Trino · Ray

The same native engine runs your Spark, Trino, and Ray jobs. Nothing to upgrade.

Iceberg · Delta Lake · Hudi

No copy and no migration: your tables stay exactly where they are.

Parquet · Avro · JSON

Flarion reads and writes your files in place, on S3, ADLS, or GCS.

Transformation · Compaction · Streaming · SQL

Nightly batch, compaction, streaming, and interactive SQL run on the same engine, at any data volume.

Forks · Odd Versions · Complex Structs · UDFs · UDAFs

We do the work to support your setup, verifying Flarion against your configuration and its edge cases in our lab.

Assess Before You Accelerate

Before you install the engine, we tell you what Flarion will do for your jobs. The assessment produces a report like this one: the expected speedup per job, and which operators are covered.
A digital report titled 'Your Flarion Report' showing workload performance. For 'your_pipeline_a,' estimated speedup is 40-50%, covering 7 of 8 operators including Scan, Filter, Explode, Shuffle, Broadcast Join, Aggregate, and Write Files. For 'your_pipeline_b,' estimated speedup is 35%, covering 5 of 7 operators including Scan, Filter, Broadcast Join, Aggregate, and Write Files. A note says Hive script transform and custom UDFs are not covered.
We send you the assessment JAR.

Read-only and auditable: no outbound calls, no permissions, jobs unchanged.

Set up and run your key workloads.

Two config lines, 6 minutes for one engineer. Then run your jobs as usual.

It rides along and takes notes.

Metadata only: your jobs’ operations and timings, plain text in your bucket.

You review, then send it over.

We simulate your workload from the file and verify performance on synthetic data.

Frequently Asked Questions

Does it work on Databricks?

Yes. Flarion deploys on Databricks with an init script.

What about UDFs and UDAFs?

They run exactly as they do today, with no code changes. Flarion doesn’t touch your bespoke code: UDF and UDAF steps run on Spark as usual, and everything around them runs native.

What happens to operations Flarion doesn’t accelerate?

They run on Spark, automatically. The rest of the job stays native and the workload completes.

How will I know what ran on Flarion vs Spark?

You can see it in Flarion’s telemetry or in your existing Spark reporting. Every stage of every run is marked native or Spark, with runtime and memory per stage.

Show more

The Latest Data Processing News & Insights

Development on Apache Spark started at Berkeley in 2009, and the first production release shipped on May 30, 2014. In the twelve years since, it has become the analytics workhorse for most of the large corporations in the world, across industries and scale, from seed-stage startups to Fortune 10 enterprises. Every year or so someone declares it old, past its peak, saddled with the JVM, and generally "legacy." And every year there is more of it. What accounts for the disconnect? In this post we'll walk through what we see across customer deployments and why we expect that in ten years there will still be a whole lot of Spark, and probably much more than there is today.

Infinite Scale

Spark scales very well. It’s not rare to see customers running workloads reading dozens of TB, while at the same time other customers process a few GB per workload. The result: for data engineering teams who don’t know how much data they’ll need to process, it’s a clean and easy decision to adopt Spark.

Network Effects

While it’s quite easy to use Spark, especially with PySpark, it’s not easy to deploy it and maintain it. But once the data platform adopts the tooling and learns how to maintain Spark, it is rarely motivated to migrate a piece of critical infrastructure to an unproven alternative, and instead is motivated to push more people to use Spark.

The Challenge of Migrating

Large companies can have thousands of jobs running at any given time, spread across the entire organization. The idea of pushing the various teams to migrate to a new platform is usually a complete non-starter. Oftentimes even gradually moving to systems like Ray is unwelcome due to the cost of maintaining multiple data platforms.

A First Class Citizen in the Data Lake

Delta Lake, Iceberg, and Hudi were each born with Spark as the reference implementation. The result is that Spark works well out-of-the-box with all three, while other systems are gradually adding support. Engineering teams want the best and most recent lakehouse technology and generally Spark supports it.

Extensibility

Spark is easy to extend without forking. Catalyst exposes optimizer rules, planning strategies, and catalog plugins. DataSource V2 lets anyone teach Spark to read a new system. User defined functions (UDFs) let teams introduce Python or Scala logic into the middle of a pipeline without leaving the framework. Plug-ins allow the introduction of new libraries into the system. The result is that the thing people would otherwise leave Spark to get, a new connector, a custom optimization, a domain-specific function library, usually shows up inside Spark instead.

The Competition

Flink is used for some streaming use cases, Ray for AI use cases, Trino for interactive SQL, DuckDB and Polars for data that fits on one machine. Data warehouses with proprietary engines are taking some share. But at this point nobody is really trying to invent a new full-fledged system to replace Spark. The competition is either specializing in a lane or building underneath it.

Improved Engines

In Spark, the underlying engine is not static. The API hasn’t changed much since DataFrames arrived, but adding Tungsten improved performance with whole stage code generation that’s close to the hardware, while query optimizations, fast paths, and new operators also make the same workload faster without code changes. Databricks added Photon, we produced Flarion, and open source brought Gluten and Comet. It’s possible to stay on Spark and get modern performance, similar to how PostgreSQL keeps getting better and adding functionality without the API changing.

Summing Up

The Spark API is likely going to be with us for a long time, but under the hood a lot is going to change. Piece by piece the engine is being replaced, and it's plausible that in ten years none of the original execution code will be left, while every job still runs and every DataFrame still looks the same. It's the Ship of Theseus, except in this version the ship gets faster with every plank. 

‍

A few weeks ago AWS shipped the Spark Upgrade Agent, an AI agent that migrates Spark jobs to Spark 4.0. You point it at a repo, it rewrites deprecated APIs, adjusts for behavioral changes, updates the build for Scala 2.13, submits the result to an EMR cluster, and iterates on failures until the job runs. It handles both Scala and PySpark, and it works the way you'd hope an agent would: plan, transform, validate, repeat. The potential payoff is large - newer Spark versions have better performance and years of accumulated bug fixes.

It looks like a good tool and data teams looking into a Spark migration should consider it, but what we’ve found is that in the enterprise, rewriting code isn’t the main impediment to upgrading Spark workloads. Spark programs are part of complex pipelines, parts of which are poorly understood or maintained, and making changes to a sensitive system is inherently risky. There’s no guarantee that the output data will actually remain the same, and data integrity is the fundamental challenge of completing such a migration.

Several companies have written in detail about major Spark upgrades, and the data integrity challenge is a recurring theme.

Slack

Slack's migration from Spark 2 to Spark 3 took about a year across 60+ EMR clusters and 40+ teams. They saw some code-level breakage: `RAND()` in join keys became an `AnalysisException`, some casts that Spark 2 tolerated started failing, the `Greatest` function handled NULLs differently than its Hive counterpart. Validation was a much larger effort. For billing pipelines, Slack required exact matches, building test tables from production data on Spark 3 and running `EXCEPT` and `COUNT` comparisons in Trino against the Spark 2 outputs, with a Python framework for digging into every discrepancy. The discrepancies weren't all bugs. Non-deterministic row ordering, timestamp variations, and genuine semantic differences between Hive and Spark implementations all produce diffs that need to be investigated. Some are noise, some are real regressions.

Uber

Uber's version of this is bigger and more instructive. They migrated from Spark 2.4 to 3.3 with over two million Spark applications running daily. The code transformation was automated with Polyglot Piranha, their structural rewrite tool. It parses the source code into an AST, matches patterns, and applies transformation rules, including inserting legacy flags like `spark.sql.legacy.allowUntypedScalaUDF` where old behavior had to be preserved. This scaled well. The problem that shaped the whole project was stated plainly: "We had over 40,000 Spark apps, so we couldn't decentralize the data validation." No staging environment, no test cases, no way to ask every team to eyeball their own outputs.

So the flagship engineering artifact of Uber's Spark upgrade wasn't actually a code migrator but Iron Dome: a shadow-testing framework which runs the migrated job against production inputs, rewrites output paths at runtime so results land in staging instead of production, puts guardrails at the Hadoop FileSystem interface so a misrouted write can't touch real data, then compares the shadow output against the production run and only marks the job migrated when they agree. 

Facebook

None of this is specific to the Spark 2-to-3 transition, or even to Spark versions. When Facebook moved Hive workloads onto Spark SQL back in 2017, they ran shadow pipelines writing to tables suffixed `_spark_shadow` so downstream jobs were never exposed, used count checks as a cheap first filter, and reached for full hash validation of outputs only reluctantly (because, as they put it, the hash validation was "sometimes even heavier than the query itself.") Funny enough, proving the new engine produced the same answer could cost more compute than producing the answer. They note that non-deterministic UDFs made validation hard, the same diff-adjudication problem Slack hit eight years later.

Takeaway

In the enterprise, migrations are rightfully considered risky projects that take time and incur risk. This is especially true in the age of AI given that the things that AI doesn’t necessarily deliver are also the riskiest parts of the migration - edge cases, data integrity, the long tail, etc. This isn’t to say that AI can’t help with building tooling for a migration, it clearly can, but usually an agent isn’t going to do the trick alone.

At Flarion, a major goal of ours is to give users the best possible performance and access to modern features without requiring a code migration. If you can get the benefits of Spark 4.2 while staying on Spark 3.4 then that’s a huge time save and reduction in risk. Of course, the data integrity problem doesn’t disappear, it’s now Flarion’s responsibility. One we’re happy to shoulder.

‍

flarion icon
Run a free assessment

See Projected Results on Your Own Jobs, Before Anything is Installed

Once deployed, every run reports the results