Skip to content

About

Code, synthetic corpus, raw traces, and results for Origin's agent-trust retrieval-poisoning experiment

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agent trust: retrieval-poisoning experiment

This repository contains the code, synthetic corpus, raw traces, and results for the experiment described in Origin's research post, Beyond Entailment: Intent, Provenance, and Agent Trust.

The experiment models an internal assistant at the fictional Auster Rille Logistics. The assistant retrieves five notices before answering an employee's question. We add false notices to the feed and measure when those notices change the answer.

What is included

  • run_exp1.py — corpus generation, embedding, retrieval, generation, and scoring
  • fixtures-auster-rille/ — the generated clean corpus metadata and the 50 target and 50 untouched facts
  • runs-auster-rille/ — every raw result cell, full JSONL traces, summary, and chart
  • PROTOCOL.md — the frozen experimental protocol
  • FINDINGS.md — results, interpretation, and limitations

The raw results include both experiment arms at 0, 1, 2, 5, and 10 inserted notices per target:

  • Naive: repeats keywords from the target question.
  • Question echo: prepends the complete target question to the false notice, following the retrieval construction used in PoisonedRAG.

Published run

The published run used:

  • Python 3
  • NumPy
  • Ollama
  • nomic-embed-text
  • gemma4:latest
  • deterministic generation (temperature=0)
  • top-five cosine retrieval
  • seed 20260827

The synthetic feed contains 2,000 filler notices, 50 target facts, and 50 untouched facts. At five question-echo notices per target, 250 notices were added to the original 2,100-document feed.

Notices per target Naive flips Question-echo flips Untouched correct, question echo
0 0/50 0/50 50/50
1 17/50 17/50 50/50
2 18/50 49/50 50/50
5 27/50 50/50 49/50
10 24/50 49/50 49/50

See FINDINGS.md for the full interpretation and limitations.

Reproduce the experiment

Install NumPy and run Ollama locally:

python -m pip install -r requirements.txt
ollama pull nomic-embed-text
ollama pull gemma4
python run_exp1.py \
  --filler 2000 \
  --targets 50 \
  --clean 50 \
  --runs-dir runs-reproduction \
  --fixtures-dir fixtures-reproduction

The run calls the Ollama API at http://127.0.0.1:11434 by default. Set OLLAMA_HOST to use another endpoint.

Results can vary when model weights, quantization, Ollama, or numerical libraries differ from the published environment. This is a small synthetic demonstration of a mechanism, not an estimate of real-world attack prevalence.

About

Code, synthetic corpus, raw traces, and results for Origin's agent-trust retrieval-poisoning experiment

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages