This repository contains the code, synthetic corpus, raw traces, and results for the experiment described in Origin's research post, Beyond Entailment: Intent, Provenance, and Agent Trust.
The experiment models an internal assistant at the fictional Auster Rille Logistics. The assistant retrieves five notices before answering an employee's question. We add false notices to the feed and measure when those notices change the answer.
run_exp1.py— corpus generation, embedding, retrieval, generation, and scoringfixtures-auster-rille/— the generated clean corpus metadata and the 50 target and 50 untouched factsruns-auster-rille/— every raw result cell, full JSONL traces, summary, and chartPROTOCOL.md— the frozen experimental protocolFINDINGS.md— results, interpretation, and limitations
The raw results include both experiment arms at 0, 1, 2, 5, and 10 inserted notices per target:
- Naive: repeats keywords from the target question.
- Question echo: prepends the complete target question to the false notice, following the retrieval construction used in PoisonedRAG.
The published run used:
- Python 3
- NumPy
- Ollama
nomic-embed-textgemma4:latest- deterministic generation (
temperature=0) - top-five cosine retrieval
- seed
20260827
The synthetic feed contains 2,000 filler notices, 50 target facts, and 50 untouched facts. At five question-echo notices per target, 250 notices were added to the original 2,100-document feed.
| Notices per target | Naive flips | Question-echo flips | Untouched correct, question echo |
|---|---|---|---|
| 0 | 0/50 | 0/50 | 50/50 |
| 1 | 17/50 | 17/50 | 50/50 |
| 2 | 18/50 | 49/50 | 50/50 |
| 5 | 27/50 | 50/50 | 49/50 |
| 10 | 24/50 | 49/50 | 49/50 |
See FINDINGS.md for the full interpretation and limitations.
Install NumPy and run Ollama locally:
python -m pip install -r requirements.txt
ollama pull nomic-embed-text
ollama pull gemma4
python run_exp1.py \
--filler 2000 \
--targets 50 \
--clean 50 \
--runs-dir runs-reproduction \
--fixtures-dir fixtures-reproductionThe run calls the Ollama API at http://127.0.0.1:11434 by default. Set OLLAMA_HOST to use another endpoint.
Results can vary when model weights, quantization, Ollama, or numerical libraries differ from the published environment. This is a small synthetic demonstration of a mechanism, not an estimate of real-world attack prevalence.