North-star objective
Make tree-sitter-analyzer the benchmark-proven leader among repository-level code understanding, navigation, and change-impact systems used by AI agents.
"No.1" is a north star, not an acceptance result. The strongest permitted claim is bounded:
On benchmark vN, for the named tools and versions, frozen repositories/questions, identical model and environment, TSA is quality/reliability non-inferior and has a statistically meaningful advantage, reproduced to E4.
Local-first operation, privacy, model/client independence, and MCP/CLI support are evaluation dimensions—not rules used to exclude otherwise relevant competitors.
Non-compensating goal tree
- G0 — Evidence integrity: immutable manifest, exact matrix, tool/repo/environment fingerprints, isolated arms, append-only attempts, checksums, deterministic replay, E0-E4 claim ladder.
- G1 — Structural/semantic correctness: definition/reference/caller/callee/import precision and recall, same-name ambiguity, dynamic dispatch, cross-language false links, exact citations.
- G2 — Real agent task effect: answer correctness/completeness, supported citations, change-task success, hallucination and wrong-edit rate under a fixed agent/model.
- G3 — Efficiency: total/cache/reasoning tokens, turns, tool calls, billed cost, warm latency, logical-cold end-to-end latency, incremental-to-answer latency, RSS and index size.
- G4 — Reliability/freshness/security: success/timeout/crash rates, exact add/modify/delete/rename refresh, zero stale rows/edges, path boundary and malicious-repository safety.
- G5 — Time to first value: clean-machine installation success, time to first correct answer, diagnostics, uninstall/recovery, Windows/macOS/Linux parity.
- G6 — Scale/language effectiveness: 10K/100K/1M LOC curves, exact indexed/error partitions, per-language conformance and real task outcomes—not README plugin counts.
- G7 — Ecosystem/adoption/reproduction: current integrations, active installs/downloads, independent contributors, second-machine blind review, public evidence and third-party E4 reproduction.
No weighted score may hide regression on another axis.
Competitor policy
- Benchmark v1 control: native file tools.
- Benchmark v1 indexed competitor: CodeGraph at a frozen version.
- GitNexus, Serena free LSP backend, CodeGraphContext, and optave/codegraph enter only after capability-boundary, install, version, and fair-adapter validation.
- Competitor inclusion rules are frozen before results. A new competitor requires a new benchmark version.
- Full IDEs/agents and SaaS products may be evaluated separately only when the model and agent scaffold are fixed and only the code-intelligence backend changes.
Hard evidence gates
Use RFC-0021 thresholds for a competitive efficiency claim:
- TSA-minus-competitor quality 95% CI lower bound > -0.10.
- Deterministic citation-location validity >= 99%.
- Success rate >= 99%.
- Any repository quality drop > 0.25 rejects the claim.
- Cost or latency ratio 95% CI upper bound <= 0.80 in at least 4/5 categories and 5/7 repositories.
- Failures, invalid trials, retries, and unfavorable experiments remain disclosed.
Multi-agent separation of powers
| Role |
Authority |
Forbidden |
| Proposer |
hypothesis, user job, objective contract |
implementation, threshold changes after results, victory claim |
| Implementer |
approved product paths in an isolated branch/worktree |
benchmark/oracle/holdout/evidence edits |
| Reviewer / Red Team |
counterexamples, mutation/error injection, security and compatibility review |
self-fixing, self-approval, lowering severity |
| Judge |
clean-checkout frozen experiment and ACCEPT/REJECT/INVALID/NOT_EVALUATED verdict |
code edits, selective reruns, gate waivers |
Human maintainer retains objective, budget, protected-path, merge, release, secret, and claim authority. Agent votes cannot establish facts.
PR value gate
CI green is necessary but insufficient. A No.1 product PR must:
- Link a frozen
NO1-xxx objective.
- Name a real user job and one primary external metric.
- Include same-environment baseline/candidate evidence.
- Meet the pre-registered minimum effect without crossing any quality/safety floor.
- Stay inside approved paths.
- Resolve independent review findings.
- Receive a clean-checkout Judge ACCEPT.
- State rollback triggers and evidence level.
Coverage, file length, project-health grade, test count, flag count, and tool count are not user value. Pure refactors do not count unless they remove a measured blocker and the associated user result improves.
Execution sequence
- Freeze grade-driven implementation while preserving read-only reporting.
- Complete NO1-001A: deterministic smoke integrity infrastructure using fixtures/dry-run only.
- Merge and freeze benchmark SHA + manifest.
- Run NO1-001B: independent three-cell Gin Smoke; it may establish only E1.
- Run a frozen 3-question x 4-repeat micro-pilot, or Gin all-questions x 4 for the RFC Pilot.
- Select one measured product gap by hard-floor violations first, then user impact x recurrence x competitor gap x verifiability.
- Implement one product hypothesis; validate on an unseen holdout.
- Warm confirmatory: 7 repositories x 21 questions x 3 arms x exactly 5 repeats = 315 cells.
- E3 second-machine independent blind review.
- E4 public evidence, at least two current indexed competitors, and third-party full reproduction.
Stop conditions
Stop and retain evidence when setup is not reproducible, a required arm is unavailable, results have already been seen before changing an oracle/threshold, an implementer crosses protected paths, the budget reaches 80% before Smoke validity, a quality/safety floor regresses, or retries/results are selectively hidden.
First objective
NO1-001A will be tracked in a separate Objective Contract issue.
North-star objective
Make tree-sitter-analyzer the benchmark-proven leader among repository-level code understanding, navigation, and change-impact systems used by AI agents.
"No.1" is a north star, not an acceptance result. The strongest permitted claim is bounded:
Local-first operation, privacy, model/client independence, and MCP/CLI support are evaluation dimensions—not rules used to exclude otherwise relevant competitors.
Non-compensating goal tree
No weighted score may hide regression on another axis.
Competitor policy
Hard evidence gates
Use RFC-0021 thresholds for a competitive efficiency claim:
Multi-agent separation of powers
Human maintainer retains objective, budget, protected-path, merge, release, secret, and claim authority. Agent votes cannot establish facts.
PR value gate
CI green is necessary but insufficient. A No.1 product PR must:
NO1-xxxobjective.Coverage, file length, project-health grade, test count, flag count, and tool count are not user value. Pure refactors do not count unless they remove a measured blocker and the associated user result improves.
Execution sequence
Stop conditions
Stop and retain evidence when setup is not reproducible, a required arm is unavailable, results have already been seen before changing an oracle/threshold, an implementer crosses protected paths, the budget reaches 80% before Smoke validity, a quality/safety floor regresses, or retries/results are selectively hidden.
First objective
NO1-001A will be tracked in a separate Objective Contract issue.