Skip to content

Goal: benchmark-proven leadership in agent code intelligence #1195

Description

@aimasteracc

North-star objective

Make tree-sitter-analyzer the benchmark-proven leader among repository-level code understanding, navigation, and change-impact systems used by AI agents.

"No.1" is a north star, not an acceptance result. The strongest permitted claim is bounded:

On benchmark vN, for the named tools and versions, frozen repositories/questions, identical model and environment, TSA is quality/reliability non-inferior and has a statistically meaningful advantage, reproduced to E4.

Local-first operation, privacy, model/client independence, and MCP/CLI support are evaluation dimensions—not rules used to exclude otherwise relevant competitors.

Non-compensating goal tree

  • G0 — Evidence integrity: immutable manifest, exact matrix, tool/repo/environment fingerprints, isolated arms, append-only attempts, checksums, deterministic replay, E0-E4 claim ladder.
  • G1 — Structural/semantic correctness: definition/reference/caller/callee/import precision and recall, same-name ambiguity, dynamic dispatch, cross-language false links, exact citations.
  • G2 — Real agent task effect: answer correctness/completeness, supported citations, change-task success, hallucination and wrong-edit rate under a fixed agent/model.
  • G3 — Efficiency: total/cache/reasoning tokens, turns, tool calls, billed cost, warm latency, logical-cold end-to-end latency, incremental-to-answer latency, RSS and index size.
  • G4 — Reliability/freshness/security: success/timeout/crash rates, exact add/modify/delete/rename refresh, zero stale rows/edges, path boundary and malicious-repository safety.
  • G5 — Time to first value: clean-machine installation success, time to first correct answer, diagnostics, uninstall/recovery, Windows/macOS/Linux parity.
  • G6 — Scale/language effectiveness: 10K/100K/1M LOC curves, exact indexed/error partitions, per-language conformance and real task outcomes—not README plugin counts.
  • G7 — Ecosystem/adoption/reproduction: current integrations, active installs/downloads, independent contributors, second-machine blind review, public evidence and third-party E4 reproduction.

No weighted score may hide regression on another axis.

Competitor policy

  • Benchmark v1 control: native file tools.
  • Benchmark v1 indexed competitor: CodeGraph at a frozen version.
  • GitNexus, Serena free LSP backend, CodeGraphContext, and optave/codegraph enter only after capability-boundary, install, version, and fair-adapter validation.
  • Competitor inclusion rules are frozen before results. A new competitor requires a new benchmark version.
  • Full IDEs/agents and SaaS products may be evaluated separately only when the model and agent scaffold are fixed and only the code-intelligence backend changes.

Hard evidence gates

Use RFC-0021 thresholds for a competitive efficiency claim:

  • TSA-minus-competitor quality 95% CI lower bound > -0.10.
  • Deterministic citation-location validity >= 99%.
  • Success rate >= 99%.
  • Any repository quality drop > 0.25 rejects the claim.
  • Cost or latency ratio 95% CI upper bound <= 0.80 in at least 4/5 categories and 5/7 repositories.
  • Failures, invalid trials, retries, and unfavorable experiments remain disclosed.

Multi-agent separation of powers

Role Authority Forbidden
Proposer hypothesis, user job, objective contract implementation, threshold changes after results, victory claim
Implementer approved product paths in an isolated branch/worktree benchmark/oracle/holdout/evidence edits
Reviewer / Red Team counterexamples, mutation/error injection, security and compatibility review self-fixing, self-approval, lowering severity
Judge clean-checkout frozen experiment and ACCEPT/REJECT/INVALID/NOT_EVALUATED verdict code edits, selective reruns, gate waivers

Human maintainer retains objective, budget, protected-path, merge, release, secret, and claim authority. Agent votes cannot establish facts.

PR value gate

CI green is necessary but insufficient. A No.1 product PR must:

  1. Link a frozen NO1-xxx objective.
  2. Name a real user job and one primary external metric.
  3. Include same-environment baseline/candidate evidence.
  4. Meet the pre-registered minimum effect without crossing any quality/safety floor.
  5. Stay inside approved paths.
  6. Resolve independent review findings.
  7. Receive a clean-checkout Judge ACCEPT.
  8. State rollback triggers and evidence level.

Coverage, file length, project-health grade, test count, flag count, and tool count are not user value. Pure refactors do not count unless they remove a measured blocker and the associated user result improves.

Execution sequence

  1. Freeze grade-driven implementation while preserving read-only reporting.
  2. Complete NO1-001A: deterministic smoke integrity infrastructure using fixtures/dry-run only.
  3. Merge and freeze benchmark SHA + manifest.
  4. Run NO1-001B: independent three-cell Gin Smoke; it may establish only E1.
  5. Run a frozen 3-question x 4-repeat micro-pilot, or Gin all-questions x 4 for the RFC Pilot.
  6. Select one measured product gap by hard-floor violations first, then user impact x recurrence x competitor gap x verifiability.
  7. Implement one product hypothesis; validate on an unseen holdout.
  8. Warm confirmatory: 7 repositories x 21 questions x 3 arms x exactly 5 repeats = 315 cells.
  9. E3 second-machine independent blind review.
  10. E4 public evidence, at least two current indexed competitors, and third-party full reproduction.

Stop conditions

Stop and retain evidence when setup is not reproducible, a required arm is unavailable, results have already been seen before changing an oracle/threshold, an implementer crosses protected paths, the budget reaches 80% before Smoke validity, a quality/safety floor regresses, or retries/results are selectively hidden.

First objective

NO1-001A will be tracked in a separate Objective Contract issue.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions