# RAG Benchmarking

> Framework-agnostic evaluation harness for RAG and agentic AI systems.

Source: https://aiexponent.com/docs/rag-benchmarking · Content verified 2026-10-04

Plug in any RAG system (LangChain, LlamaIndex, or custom) and benchmark it against classic and agentic-era metrics. Faithfulness, answer relevancy, retrieval precision, and four agentic metrics for multi-step agents. Measured faithfulness of 0.958 on the 50-sample golden dataset.

- EU AI Act: Article 15 (Accuracy Requirements)
- PyPI package: `rag-benchmarking` v1.0.2
- Source: https://github.com/aiexponent/rag-benchmarking
- Licence: Apache License 2.0

## Quick start

```bash
pip install rag-benchmarking
```

```python
from rag_benchmarking import RagEval

client = RagEval(api_url="http://localhost:5001", api_key="your-key")

# Works with LangChain
result = my_chain.invoke({"query": "What is RAG?"})
sample = RagEval.from_langchain(result)

# Or any dict with question / contexts / answer
sample = {
    "question": "What is RAG?",
    "contexts": ["RAG stands for Retrieval-Augmented Generation."],
    "answer": "RAG combines retrieval with LLM generation.",
}

report = client.evaluate([sample], metrics=["faithfulness", "answer_relevancy"])
print(report["metrics"])
# {"faithfulness": 0.958, "answer_relevancy": 0.810}
```

## For coding agents

Install rag-benchmarking from PyPI into a virtual environment (Python 3.11 or newer). Retrieval metrics need no API key: import `EvalSample` and `RunConfig` from `rag_benchmarking.harness.schemas` and `EvaluationRunner` from `rag_benchmarking.harness.runner`, build samples with `retrieved_doc_ids` and `relevant_doc_ids`, then call `EvaluationRunner(RunConfig(metric_group="retrieval")).evaluate(samples)`. The LLM-judge metrics need GEMINI_API_KEY; on a fresh install of 1.0.2 they fail to import (github.com/aiexponent/rag-benchmarking/issues/35). Docs: https://aiexponent.com/docs/rag-benchmarking.md

## Benchmark results

Measured on the 50-sample golden dataset using `gemini-2.5-flash` as judge at `temperature=0.0`.

| Metric | Score | Note |
| --- | --- | --- |
| faithfulness | 96% | Excellent |
| answer_relevancy | 81% | Good |

## Features

- Framework-agnostic: works with LangChain, LlamaIndex, or any custom RAG system
- Classic metrics: faithfulness, answer relevancy, context precision/recall
- Retrieval metrics: Precision@K, Recall@K, MRR, NDCG
- Agentic metrics: agent faithfulness, tool call accuracy, source attribution, retrieval necessity
- REST API + Python SDK with LangChain and LlamaIndex adapters
- Run history with comparison across configurations

## Regulatory foundation

### Article 15: Accuracy, robustness and cybersecurity

Status: UPCOMING. Applies from 2 Dec 2027 · deferred (statutory date before the Digital Omnibus: 2026-08-02).

> 1. High-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness, and cybersecurity, and that they perform consistently in those respects throughout their lifecycle.
>
> 3. The levels of accuracy and the relevant accuracy metrics of high-risk AI systems shall be declared in the accompanying instructions of use.
>
> 4. High-risk AI systems shall be as resilient as possible regarding errors, faults or inconsistencies that may occur within the system or the environment in which the system operates, in particular due to their interaction with natural persons or other systems. Technical and organisational measures shall be taken in this regard. The robustness of high-risk AI systems may be achieved through technical redundancy solutions, which may include backup or fail-safe plans.

Paragraphs quoted: 15(1), 15(3), 15(4). Source: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689 (retrieved 2026-10-04).

Penalty: Up to €15M or 3% of global annual turnover, whichever is higher (Article 99(4), https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689).

Article 15 applies from 2 December 2027 for high-risk AI systems under Annex III. The original 2 August 2026 date was moved by the Digital Omnibus, Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force since 27 July 2026. Providers must declare accuracy metrics in the instructions for use and demonstrate consistent performance across the lifecycle. Non-compliance via the Article 16 provider-obligation chain is sanctionable up to €15M or 3% of global annual turnover under Article 99(4). For RAG-based high-risk systems, "appropriate accuracy" is not a self-asserted figure. It is a metric declared on the label and defensible against post-market evidence.

How RAG Benchmarking addresses this:

- 15(1): Reproducible accuracy benchmarks for RAG pipelines (retrieval recall, answer faithfulness, citation precision) with versioned eval sets
- 15(3): Generates the accuracy-metrics block for the Article 13 instructions for use, with confidence intervals and eval-set provenance
- 15(4): Robustness suite: input perturbations, noisy-context, adversarial-passage, and OOD query stress tests with pass/fail thresholds
- 15(4): Lifecycle drift monitoring: replays the declared eval set against the live system on a schedule and alerts on metric regression

## FAQ

### What does EU AI Act Article 15 require?

High-risk AI systems must achieve appropriate accuracy, robustness, and cybersecurity throughout their lifecycle. Accuracy metrics must be declared in the instructions for use (Article 15(3)), and the system must be resilient to errors, faults, and inconsistencies (Article 15(4)). Source: Regulation (EU) 2024/1689 Article 15(1), 15(3), 15(4).

### When does Article 15 become enforceable?

Article 15 obligations apply from 2 December 2027 for high-risk AI systems under Annex III, and from 2 August 2028 for systems under Annex I. The Digital Omnibus, Regulation (EU) 2026/1744, moved the original 2 August 2026 date. Parliament adopted it on 16 June 2026 and the Council on 29 June 2026; it was published in the Official Journal on 24 July 2026. Source: Regulation (EU) 2024/1689 Article 113, as amended by Regulation (EU) 2026/1744.

### Does RAG Benchmarking cover the cybersecurity leg of Article 15?

No. RAG Benchmarking covers the accuracy and robustness legs: faithfulness, retrieval precision, agentic metrics, and adversarial-passage stress tests. The cybersecurity leg (prompt injection resistance, jailbreak defence, model integrity) needs a runtime AI security control. Pair the two to cover both legs of Article 15.

### What metrics does RAG Benchmarking measure?

Classic metrics (faithfulness, answer relevancy, context precision/recall), retrieval metrics (Precision@K, Recall@K, MRR, NDCG), and four agentic metrics (agent faithfulness, tool-call accuracy, source attribution, retrieval necessity).

### Is RAG Benchmarking framework-agnostic?

Yes. RAG Benchmarking works with LangChain, LlamaIndex, or any custom RAG system that returns a sample with `question`, `contexts`, and `answer` fields. SDK adapters for LangChain and LlamaIndex are included; custom integrations use the JSONL schema directly.

### What is the measured faithfulness on the golden dataset?

0.958 on the published 50-sample golden dataset (rated "Excellent"), with 0.810 answer relevancy ("Good"). These are the actual numbers from the v1.0.0 release benchmark, not aspirational targets.

### Can I bring my own evaluation dataset?

Yes. RAG Benchmarking accepts custom datasets in JSONL format with the expected schema. The bundled golden dataset is English-only; multilingual evaluation is not supported in v1.0.

### Is RAG Benchmarking free?

Yes. Apache 2.0 licensed. The harness itself runs locally; LLM-as-judge metrics depend on whichever judge model you configure (which may have its own usage cost).

### What is the penalty for Article 15 non-compliance?

Up to €15M or 3% of global annual turnover, whichever is higher, under Article 99(4). The provider-obligation chain via Article 16 routes Article 15 failures through this penalty band.

### How does drift monitoring work?

You declare an evaluation set version and a metric threshold. RAG Benchmarking replays the eval set against the live system on a schedule and alerts on metric regression, supporting the lifecycle-consistent-performance requirement of Article 15(1).

## Known limitations

- Benchmark datasets are English-only; no multilingual evaluation support.
- Custom dataset integration requires manual formatting to the expected JSONL schema.
- Accuracy metrics only; latency and throughput are not measured.
- LLM-as-judge metrics depend on the configured judge model quality.
- Rate limiting is in-memory and resets on server restart.

## Contributing

Issues: https://github.com/aiexponent/rag-benchmarking/issues · Contributing guide: https://github.com/aiexponent/rag-benchmarking/blob/main/CONTRIBUTING.md

---

Not legal advice. Not a notified body. The tools produce evidence, not conformity assessment.
All docs as Markdown: https://aiexponent.com/llms.txt · Guide for coding agents: https://aiexponent.com/agents.md
