A cost-aware retrieval intelligence layer that determines when semantic search is actually necessary.
ARI is an adaptive retrieval middleware layer for AI agents, RAG systems, and enterprise search applications. It dynamically selects the least expensive retrieval strategy capable of satisfying an information need, escalating to deeper retrieval only when evidence is insufficient.
The project combines a production-oriented software platform with a reproducible research benchmark, allowing the community to improve individual components without changing the core orchestration layer.
- Adaptive Retrieval Intelligence (ARI)
- Table of Contents
- Vision
- Problem
- Core Hypothesis
- Objectives
- Research Questions
- Reference Architecture
- Core Modules
- Retrieval Strategy
- Product & API Vision
- Repository Structure
- Community Workstreams
- Research & Evaluation Discipline
- First Prototype
- 16-Week Roadmap
- v0.1 Success Criteria
- Risk Management
- Open-Source Philosophy
- License
- Status
Modern RAG systems frequently default to semantic retrieval even when simpler approaches such as exact matching, BM25, or fuzzy search would have been sufficient.
ARI introduces an adaptive retrieval decision layer between the application and the retrieval infrastructure.
Its goal is simple:
Use the minimum retrieval intelligence necessary to reliably answer the information need.
Instead of treating semantic search as the default, ARI treats it as an escalation mechanism.
This enables retrieval systems to optimize simultaneously for:
- Retrieval quality
- Latency
- Compute consumption
- Infrastructure cost
- API usage
- Semantic-search utilization
- Repeat-query efficiency
A typical modern retrieval pipeline may involve:
- Query embedding
- Vector database search
- Candidate generation
- Hybrid retrieval
- Reranking
- LLM-based evaluation or synthesis
These operations can introduce significant compute, latency, and API costs.
However, not every query requires the same level of retrieval intelligence.
For example:
- An exact identifier may be resolved through exact matching.
- A structured business query may perform well with BM25.
- A typo-heavy query may benefit from fuzzy retrieval.
- A conceptual query may require semantic retrieval.
- An ambiguous query may require hybrid retrieval and reranking.
ARI attempts to determine which strategy is appropriate for each query rather than applying the most expensive strategy universally.
Semantic search should be an escalation mechanism rather than the default retrieval mechanism.
ARI tests whether adaptive routing can reduce retrieval cost and latency while maintaining an agreed quality threshold.
ARI aims to:
- Reduce unnecessary semantic retrieval.
- Reduce embedding generation and vector-search activity.
- Reduce unnecessary LLM calls.
- Reduce retrieval latency.
- Reduce infrastructure and external API costs.
- Maintain an agreed retrieval-quality threshold.
- Learn which retrieval strategies work best for different query classes and corpora.
- Reuse previous retrieval work through caching.
- Reduce repeated computation through query-delta processing.
- Provide measurable and reproducible evidence for every optimization.
The project investigates the following questions:
- Can query characteristics predict whether semantic retrieval is necessary?
- Can lexical retrieval satisfy a larger proportion of enterprise queries than systems that default to semantic search?
- Can confidence-driven escalation maintain retrieval quality while reducing latency and computational cost?
- How accurately can a Semantic Necessity Score (SNS) predict the value of semantic retrieval?
- Can Retrieval Strategy Memory improve routing without requiring full model retraining?
- How much embedding, vector-search, and LLM activity can be avoided?
- Can retrieval caching and query-delta processing reduce repeated computation?
- How does adaptive retrieval behave across different corpora, languages, and workload distributions?
┌──────────────────────┐
│ AI Agent / RAG │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Query Intelligence │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Semantic Necessity │
│ Engine (SNS) │
└──────────┬───────────┘
│
▼
┌──────────────────────┐
│ Retrieval Router │
└──────────┬───────────┘
│
┌────────────────┼────────────────┐
│ │ │
▼ ▼ ▼
┌───────────┐ ┌───────────┐ ┌──────────────┐
│ Lexical │ │ Hybrid │ │ Semantic │
│ Exact │ │ Retrieval │ │ Vector │
│ BM25 │ │ │ │ Reranking │
│ Fuzzy │ │ │ │ │
└─────┬─────┘ └─────┬─────┘ └──────┬───────┘
│ │ │
└────────────────┼────────────────┘
▼
┌──────────────────────┐
│ Confidence Engine │
└──────────┬───────────┘
│
┌─────────┴─────────┐
│ │
▼ ▼
Evidence sufficient Insufficient
│ │
▼ ▼
Return Escalation Engine
│
▼
Deeper Retrieval
│
▼
Strategy Memory
│
▼
Cost Intelligence
│
▼
Cache
Responsible for understanding the incoming information need.
Capabilities include:
- Query normalization
- Entity extraction
- Query classification
- Intent detection
- Complexity estimation
- Query-feature generation
The Semantic Necessity Engine calculates a Semantic Necessity Score (SNS) between 0 and 1.
The engine should return both:
- A numeric score
- An auditable explanation object describing the factors behind the decision
Example:
{
"score": 0.72,
"reason": "Conceptual query with low lexical overlap",
"signals": [
"low_exact_match_probability",
"high_semantic_intent",
"ambiguous_terminology"
]
}Selects the least expensive retrieval strategy appropriate for the query.
Possible strategies include:
- Exact retrieval
- BM25
- Fuzzy retrieval
- Query expansion
- Hybrid retrieval
- Semantic/vector retrieval
- Reranking
Provides low-cost retrieval primitives:
- Exact matching
- BM25
- Fuzzy retrieval
- Query expansion
Combines lexical and semantic candidates using configurable fusion strategies.
Provides deeper retrieval capabilities:
- Embedding generation
- Vector retrieval
- Optional reranking
- Semantic candidate generation
Determines whether the retrieved evidence is sufficient to satisfy the query.
The confidence decision becomes the primary trigger for escalation.
Moves progressively from cheaper retrieval strategies toward deeper retrieval only when required.
A typical path may be:
Exact
↓
BM25
↓
Fuzzy / Expanded
↓
Hybrid
↓
Semantic
↓
Reranking
Not every query needs to traverse the entire pipeline.
Stores historical retrieval performance by:
- Query class
- Corpus
- Retrieval strategy
- Workload
- Success/failure outcome
- Escalation path
The objective is to allow ARI to learn from previous retrieval decisions without requiring full model retraining.
Tracks and estimates:
- CPU usage
- GPU usage
- Embedding calls
- Vector searches
- LLM calls
- Token consumption
- External API calls
- Latency
- Retries
- Estimated cost
- Actual cost
Supports:
- Exact-query caching
- Normalized-query caching
- Semantic retrieval-result caching
- Document/result reuse
- Query-delta processing
- Freshness metadata
Measures:
- Retrieval quality
- Task success
- Latency
- Cost
- Escalation behaviour
- Compute avoidance
ARI should optimize for minimum sufficient retrieval intelligence.
A simplified decision flow is:
Incoming Query
│
▼
Query Intelligence
│
▼
SNS Score
│
┌────────────┴────────────┐
│ │
Low necessity High necessity
│ │
▼ ▼
Lexical retrieval Hybrid / Semantic
│ │
└────────────┬────────────┘
▼
Confidence Check
│
┌────────┴────────┐
│ │
Sufficient Insufficient
│ │
▼ ▼
Return Escalate
The key design principle is that retrieval depth should be evidence-driven rather than predetermined.
ARI is intended to operate as retrieval middleware for AI applications.
POST /retrieve
GET /strategy/{query_id}
GET /metrics
POST /feedback
POST /index
GET /cache/{key}A typical request lifecycle:
Client
│
▼
POST /retrieve
│
▼
Query Analysis
│
▼
Strategy Selection
│
▼
Retrieval
│
▼
Confidence Evaluation
│
├── Sufficient ───────► Return Results
│
└── Insufficient ─────► Escalate
│
▼
Deeper Retrieval
│
▼
Return Results
The response should expose retrieval metadata so that routing decisions remain measurable and auditable.
core/ # Router, interfaces, confidence, configuration
query/ # Classification, normalization, feature extraction
lexical/ # Exact, BM25, fuzzy retrieval, query expansion
semantic/ # Embeddings, vector retrieval, reranking
hybrid/ # Fusion and adaptive fusion
memory/ # Strategy memory and feedback
cost/ # Cost models and optimization
cache/ # Retrieval caching and query-delta processing
integrations/ # RAG and agent adapters
benchmarks/ # Datasets, runners and reports
evaluation/ # Quality, latency and cost metrics
examples/ # Reference applications
docs/ # Architecture, research and contribution guides
The project is intentionally modular so contributors can focus on individual research or engineering areas.
- Query classification and intent detection
- Lexical retrieval optimization
- Semantic Necessity Score
- Confidence and escalation algorithms
- Retrieval Strategy Memory
- Cost and energy models
- Semantic caching
- Multilingual retrieval
- RAG and agent integrations
- Benchmark datasets and evaluation
Improve one component. Measure the result. Contribute the evidence.
ARI is both an engineering project and a research programme.
Every optimization should be evaluated against meaningful baselines.
Compare ARI against:
| System | Retrieval Strategy |
|---|---|
| Baseline A | Always lexical / BM25 |
| Baseline B | Always semantic / vector |
| Baseline C | Fixed hybrid |
| ARI | Adaptive routing |
Measure:
- Precision
- Recall
- Relevance
- Downstream task success
- Median latency
- P95 latency
- P99 latency
- Embedding calls
- Vector searches
- LLM calls
- Token consumption
- Infrastructure/API cost
- Semantic escalation rate
- Compute avoidance rate
The central result should quantify:
How much cost and latency ARI removes while preserving an agreed quality threshold.
These results must be experimentally measured, reproducible, and reported against fixed baselines.
The first prototype should remain deliberately small and measurable.
- 5,000–10,000 queries
- BM25 retrieval
- Hybrid retrieval
- Semantic retrieval
- Basic SNS model
- Confidence-based escalation
- End-to-end cost benchmark
- End-to-end latency benchmark
Can ARI reduce semantic retrieval usage and retrieval cost while maintaining acceptable retrieval and downstream task quality?
The first release should prioritize experimental clarity over feature breadth.
| Phase | Weeks | Major Output |
|---|---|---|
| Foundation | 1–2 | Architecture + benchmark framework |
| Lexical | 3–4 | Exact / BM25 / Fuzzy retrieval |
| SNS | 5–6 | Semantic Necessity Engine |
| Escalation | 7–8 | Adaptive Retrieval Router |
| Memory | 9–10 | Retrieval Strategy Memory |
| Cost | 11 | Cost Intelligence |
| Cache | 12 | Caching + query-delta processing |
| Integration | 13 | RAG / agent adapters |
| Evaluation | 14 | Benchmark report |
| Release | 15–16 | Open-source v0.1 |
The v0.1 release should satisfy the following criteria:
- ARI exposes a working retrieval API.
- Queries can be routed across multiple retrieval strategies.
- SNS produces an auditable routing signal.
- Escalation is driven by retrieval confidence.
- Strategy Memory records historical retrieval performance.
- Cost and latency are measured per query.
- ARI is benchmarked against lexical, semantic, and hybrid baselines.
- At least one RAG or agent integration is demonstrated.
- Experiments are reproducible from the repository.
- Results document both quality and efficiency trade-offs.
| Risk | Mitigation |
|---|---|
| Routing errors | Confidence thresholds and fallback strategies |
| Poor benchmark quality | Expert-reviewed evaluation sets |
| Corpus dependence | Held-out corpora and query classes |
| Overfitting | Reproducible benchmark splits and independent evaluation |
| Stale cache results | Freshness metadata and invalidation policies |
| Cost-model errors | Predicted-vs-actual cost validation |
| Community fragmentation | Stable interfaces and modular component boundaries |
ARI is designed as a collaborative research and engineering project.
Contributions are welcome across:
- Code
- Research
- Datasets
- Benchmark cases
- Documentation
- Integrations
- Examples
- Reviews
- Evaluation methodology
- Performance improvements
The project values measurable improvements over assumptions.
Build one component. Measure it. Share the result. Improve the ecosystem.
Choose and publish an explicit OSI-approved open-source license in LICENSE before the first public release.
The repository should also document:
- Contribution guidelines
- Code of conduct
- Development setup
- Benchmark reproduction instructions
- Citation information
- Security reporting process
Project stage: Research + prototype
Target release: v0.1
Primary objective: Determine whether adaptive retrieval can materially reduce semantic retrieval usage, latency, and cost without compromising retrieval quality.