# Retrieval-Augmented Generation

## Ground AI in enterprise data

Build RAG pipelines that combine live, structured, and unstructured data with SQL and hybrid search. Retrieve, rank, and feed real-time context directly to your models.

[Get a demo](https://meetings.hubspot.com/lukekim/talk-to-sales?uuid=836fd7be-a95e-4cee-b0cb-044fd8ea52a4&utm_campaign=22505145-Website+Demo+Requests&utm_source=homepageheader&utm_medium=website&utm_content=demo+request) [View the docs](https://spiceai.org/docs/use-cases/rag/applications)

### Do more with your data

100x

up to 100x faster queries

80%

up to 80% cost savings on data lakehouse spend

2x

Increase in data reliability

### RAG breaks down without unified, real-time data

Fragmented RAG pipelines force developers to manage multiple search engines, connectors, and model APIs. Models are grounded in incomplete or outdated data, leading to hallucinations, inconsistencies, and production risks.

### Build data-grounded AI faster

Deliver accurate, context-rich results by combining SQL federation, hybrid search, and LLM inference in one governed runtime.

#### Federate structured data

Query operational and analytical data across databases, object stores, and APIs using standard SQL. All in real time with zero ETL.

[Learn more about Spice federation](/content/platform/sql-federation-acceleration/index.html)

#### Search and embed unstructured data

Create embeddings from text, documents, or logs using local or hosted models. Run vector similarity search directly inside Spice to find relevant context, semantic matches, and related insights instantly.

[Explore hybrid SQL search](/content/platform/hybrid-sql-search/index.html)

#### Retrieve and augment intelligently

Blend structured SQL results and semantic search results in one query using hybrid search. Deliver precise, context-rich data to your LLMs without manual data engineering or pipeline maintenance.

[Learn about real-time hybrid search using RRF](/content/blog/real-time-hybrid-search-using-rrf/index.html)

#### Generate insights with the AI Gateway

Invoke and run models like OpenAI, Anthropic, or local LLMs directly in SQL queries. Feed retrieved context into the model and generate accurate, compliant, and contextual results within the same SQL workflow.

[Explore LLM inference and AI model serving](/content/platform/llm-inference/index.html)

### Why choose Spice for RAG

Spice unifies data retrieval, semantic search, and AI generation in a single, high-performance runtime-no pipelines, no orchestration, no drift.

Real-Time Federation

Query all your sources with federated SQL. No data movement or batch syncs.

Built-in Vector Search

Embed and retrieve semantic context alongside SQL filters and full-text search.

SQL LLM Inference

Call, prompt, and evaluate models in SQL via the AI() SQL function.

Hybrid Ranking

Blend multiple result sets with Reciprocal Rank Fusion for per-query weighting and tunable relevance.

Governed & Secure

Enterprise-grade access control ensures compliance and auditability.

Deployment Flexibility

Run Spice anywhere: as a sidecar, microservice, cluster, or on the managed Spice Cloud Platform.

### Trusted by teams building intelligent applications

Run data-intensive workloads on a high-performance engine trusted by teams building real-time systems at scale.

#### “Partnering with Spice AI has transformed how NRC Health delivers AI-driven insights. By unifying siloed data across systems, we accelerated AI feature development, reducing time-to-market from months to weeks - and sometimes days. With predictable costs and faster innovation, Spice isn't just solving some of our data and AI challenges - it's helping us redefine personalized healthcare.”  
Tim Ottersburg  
VP of Technology, NRC Health

#### “Spice AI grounds AI in our actual data, using SQL queries across many data sources. This brings accuracy to probabilistic AI systems, which are very prone to hallucinations.”  
Rachel Wong  
CTO, Basis Set

### Build a scalable RAG app

Guides and examples to learn more about building RAG applications with Spice.

[Hybrid Search Docs](/content/vc-ap-d1befa/_next/image?url=%2Fwebsite-assets%2Fmedia%2F2025%2F12%2FDocs-jpeg.jpeg&w=2048&q=75&dpl=dpl_2oWJeni9AeWgJWMFPcdvgvgtpJYJ/index.html)

**Hybrid Search Docs** 
Spice provides robust search capabilities enabling developers to query datasets beyond traditional SQL, including semantic (vector-based) search, full-text keyword search, and hybrid search methods.

[True Hybrid Search: Vector, Full-Text, and SQL in One Runtime](/content/blog/true-hybrid-search/index.html)

**True Hybrid Search: Vector, Full-Text, and SQL in One Runtime** 
TL;DR  Show Me the (Data)! It’s well established (and maybe even trite) to say that enterprises are going all-in on artificial intelligence, with more than $40 billion directed toward generative AI projects in recent years. The initial results have been underwhelming. A recent study from the Massachusetts Institute of Technology's NANDA initiative concluded that despite the enormous allocation of

[Real-Time Hybrid Search Using RRF: A Hands-On Guide with Spice](/content/blog/real-time-hybrid-search-using-rrf/index.html)

**Real-Time Hybrid Search Using RRF: A Hands-On Guide with Spice** 
Surfacing relevant answers to searches across datasets has historically meant navigating significant tradeoffs. Keyword (or lexical) search is fast, cheap, and commoditized, but limited by the constraints of exact matching. Vector (or semantic) search captures nuance and intent, but can be slower, harder to debug, and expensive to run at scale. Combining both usually entails standing up multiple engines.

### See Spice in action

Walk through your use case with an engineer and see how Spice handles federation, acceleration, and AI integration for production workloads.

[Talk to an engineer](https://meetings.hubspot.com/lukekim/talk-to-sales)
