Spice 2.0 is Now Available: Real-Time Analytical Query on Operational Data, Without ETL | Spice AI | Spice AI

Spice 2.0: Real-Time Analytical Query on Operational Data, Without ETL

Spice AI

Releases

Spice 2.0

Luke Kim

Founder and CEO of Spice AIJuly 9, 2026

AI agents that need real-time data are driving new, demanding, analytical workloads on operational databases.

However, analytical queries on stores like MySQL, PostgreSQL, and MongoDB risk disrupting mission-critical operations with heavy execution, complex RLS-policies, and data-leakage. ETL pipelines that copy data from operational to analytical systems are expensive, costly to operate, and are not real-time, typically with stale data in the hours or even days.

With Spice 2.0, organizations can add sandboxed analytics replicas alongside their operational databases in minutes with sub-second query and real-time freshness, using high-throughput replication.

And when the data outgrows one node, deploy petabyte-scale compute with confidence. Multi-node, multi-active, highly available distributed query built on Apache Ballista is now generally available.

What's New in Spice 2.x

High-Throughput Change-Data-Capture (CDC) Replication. Bolt high-performance analytics-ready replicas onto live operational databases. Spice replicates directly from the PostgreSQL WAL, MySQL binlog, and MongoDB oplog without pipelines or query load on production. It's incrementally adoptable: start with a single table and be running analytical queries on operational data in minutes with 2-second end-to-end freshness under continuous ingest. v2.1 adds an in-memory CDC tier and a dedicated compaction runtime that cut replication lag on high-volume workloads, plus shared PostgreSQL replication slots across changes-mode datasets.

Multi-Node Distributed Compute. Petabyte-scale compute built on Apache Ballista is now generally available. Object-store native and highly available, with multi-active schedulers and no single point of failure. Three executors run TPC-H SF100 2.9x faster than one node. v2.1 distributes Iceberg catalog table scans and broadcast-joins small dimension tables, with shared scheduler job state and failover.

Spice Cayenne. The premier Spice acceleration engine built on Vortex is now generally available: 1.5x faster than DuckDB with 3x less steady-state memory, and 26x faster than Spice 1.x on TPC-DS SF100 - now with atomic WAL-staged writes, high-throughput ingestion, MERGE INTO, and SQL-defined partitioning. v2.1 adds experimental adaptive self-tuning that adapts Cayenne to hardware, schema, and live workload.

Spice Kubernetes Operator. Full lifecycle management and control for running Spice at scale on Kubernetes with blue/green deployments, instant-rollback, and data-aware routing.

Enterprise Security & Policy. OIDC authentication, a Cedar policy engine, PII masking, and mTLS, all enforced before data reaches the agent.

And More! Searchable Tool Registry, SQL & WASM user-defined functions (UDFs), DataFusion v54, and new data connectors including Elasticsearch and Azure Cosmos DB.

Benchmarks: 1.5 - 1.8x faster than Spice 1.x on TPC-H SF100 with up to 42% less memory · 26x faster on TPC-DS SF100 with Cayenne · ~170x faster CDC ingest than 1.x · 2.9x faster distributed query from one node to three · 2-second CDC freshness · 1,046 QPH of CH-BenCHmark analytics at SF1000 under a 266,000+ tpmC live transactional load - full results below.

Performance

Spice 2.0 numbers are measured on release-candidate builds; Spice 1.x numbers on the final 1.x release (v1.11.6) - identical harness, hardware, and benchmark specs throughout. Cluster benchmarks ran on i3.4xlarge nodes (16 vCPU, 122 GB RAM) in AWS; single-node benchmarks ran under a 256 GB memory limit.

Analytics on live operational data: CH-BenCHmark

CH-BenCHmark is the classic HTAP (hybrid transactional/analytical processing) benchmark: it runs the TPC-C transactional workload and TPC-H-style analytical queries concurrently, against the same data. Benchmarks like TPC-H measure analytics on data at rest; CH-BenCHmark measures the scenario Spice 2.0 is built for - analytical queries that stay fast and correct while transactions continuously change the data underneath them.

The benchmark configuration mirrors the 2.0 architecture end-to-end: PostgreSQL serves the TPC-C transactional workload, a single Spice node replicates committed changes via CDC into Cayenne acceleration, and analytical queries run against the Spice replica - production never sees the analytical load. Scale factor 1000: 1,000 warehouses and 300M+ rows, with a 600-second measurement window on a single 64-core node.

Metric Result
Source bootstrap (300M-row order_line) ~9 minutes via native PostgreSQL logical replication - ~566K rows/s
Bootstrap ingest rate vs. Spice 1.x ~170x faster than the 1.x Debezium-based path
Transactional throughput (PostgreSQL) 266,861 tpmC - 6.09M transactions in ~10 minutes per node
Analytical throughput, concurrent with ingest 1,046 QPH per node

PostgreSQL sustained roughly 10,000 transactions per second for the full window while Spice served the entire analytical workload from the replica.

Operational freshness under load

Spicebench measures the full operational loop on TPC-H SF10: continuous ingest, concurrent query load, and checkpoint validation until queries return exactly-correct results. End-to-end freshness - from ingest to exactly-correct results - is single-digit seconds in both modes: 2.0 s for CDC (inserts + updates + deletes) and 7 s for append streams.

Spice 2.0 vs. Spice 1.x: same hardware, same data

Benchmark · accelerator Spice 1.x Spice 2.0 Spice 2.0 is
TPC-H SF100 · DuckDB 253.0 s 138.3 s 1.8x faster
TPC-H SF100 · Cayenne 133.6 s 88.3 s 1.5x faster
TPC-DS SF100 · DuckDB 108.6 s 93.9 s 16% faster
TPC-DS SF100 · Cayenne 4,196 s 157.6 s 26x faster
Peak memory · TPC-H SF100 · DuckDB 70.7 GB 40.7 GB 42% less

Cluster-Sidecar Architecture

The new multi-cluster support enables new workload-optimized architectures. The cluster-sidecar architecture combines a Spice multi-node cluster for scale and lightweight sidecars for locality, so agentic workloads get both fast, low-latency query of hot data and fast, distributed query across petabyte scale data lakes.

In this pattern, Single-node Spice instances run co-located with applications, materializing only the working set that application or AI agent requires. When a sidecar needs to reach beyond its local working set or run heavy long-running queries, it delegates to the cluster.

The tiers collaborate rather than just stack. The multi-node cluster handles ingestion, replication, shared accelerations, while each sidecar stays light and specific to its agent or dashboard, serving and caching the working set it needs.

The sidecar acts as a physically isolated sandbox of working data, while Spice manages refreshing data and delegating queries to the cluster replicas instead of the production database, so analytical queries and agents never contend with the operational workload.

Spice Cayenne (GA)

Spice Cayenne was introduced in preview in v1.9 as Spice's next-generation data accelerator and reaches general availability in 2.0.

It's built on Vortex, an open-source columnar format (Linux Foundation, Apache-licensed) designed for the access patterns of agents: random access, point lookups, concurrent readers, and streaming updates (unlike Parquet's batch analytics focus). Cayenne pairs a Vortex data layer with an embedded metadata engine, which supports terabyte-scale workloads with significantly lower memory requirements than DuckDB.

With 2.0, Cayenne now supports high-throughput ingestion of writes, changes, and deletes beyond its append-optimized initial release. Writes are staged through a write-ahead log and commit atomically with full ACID semantics. Small writes are absorbed by a low-latency inline/mem-tier and are immediately queryable.

For CDC, primary-key DELETEs that identify keys directly skip the table scan, updates use merge-on-read position deletes instead of rewriting the table, and a dedicated compaction runtime keeps that background work off the query and ingest. MERGE INTO plus SQL-defined (PARTITION BY) and composite partitioning round out the SQL surface.

Production characteristics:

Spice Kubernetes Operator

The Spice Kubernetes Operator v1.0 is now generally available to enterprise customers. It manages the full data-aware deployment lifecycle, on cloud or on-premises, for running Spice at scale on Kubernetes.

Two Custom Resource Definitions (CRDs) are included supporting multi-node clusters and single-node deployments. The SpicepodCluster CRD manages scheduler and executor nodes, mTLS certificate provisioning, rolling upgrades, and automatic failover. The SpicepodSet CRD manages single-node deployments and sidecars: injecting Spice instances into application pods via annotations, scaling them horizontally, and handling persistent storage when required.

Enterprise Security & Control

Spice 2.0 includes the security and controls demanded by enterprises and new agentic workloads.

The Spice platform is secure-by-default with mTLS, OIDC authentication, and RBAC and ABAC authorization. Agent working sets of data are declaratively defined so each sandboxed Spice instance is only provisioned with data that any specific agent should access. Fine-grained policy to the row and column level can be defined and enforced by the Cedar policy engine, especially useful for defining specific data LLM tools or UDFs can access. Cedar policy is integrated and enforced in the core DataFusion query engine, with no ability to circumvent via SQL.

The Spice Cloud Hybrid Model

Spice Cloud is a managed Spice service that operates cloud-hosted multi-node Spice clusters, including high-availability distributed query, Cayenne acceleration, search, and AI inference for you.

In a hybrid cluster-sidecar deployment, Spice Cloud manages the multi-node cluster while application sidecars can run in your own environment alongside your applications and agents. Heavy compute is delegated to the fully managed infrastructure. Latency-sensitive sidecar instances run wherever your applications and agents live, in your Kubernetes clusters, VPCs, on-premises data centers, or edge, and serve hot data locally at sub-second latency over mTLS.

What's next

We're already working on the next chapter of Spice: BYOC (bring-your-own-cloud), distributed search across multi-node clusters, write-back acceleration with full DML, an Iceberg-REST compatible Cayenne Catalog, webhooks and event-driven actions, and much more. The roadmap is public and community-driven.

The demands on data infrastructure keep growing as apps and agents demand real-time operational, analytical, streaming, and service data. From our founding in 2021 to the future, we're building Spice as the best data platform to power the next-generation of intelligent, AI-driven apps and agents.

Add Spice to your operational data. Analytical query with no ETL. It's open source, portable, scalable, and fast. Welcome to Spice 2.0.

If you're an architect or technical leader evaluating data infrastructure for AI agents in production, we'd love to talk. Mention this post when you reach out to hey@spice.ai, and the first 15 teams will receive a dedicated architecture workshop with our engineering team.