Skip to content

[Spike] Establish a defensible synthetic-label and score-governance process #73

Description

@knytcomics-ui

Category

Spike

Question

What evidence and review process should determine labels and static scores while preserving a deterministic synthetic dataset?

Context

Labels are synthetic and notes-based; scores are static, while research plans accuracy and false-positive measurements.

Why This Matters

Evaluation is only meaningful when labels and scores are explainable, repeatable, and protected from silent bias.

Areas to Investigate

Expert review, adversarial clean controls, score calibration, reviewer agreement, provenance, exceptions, and synthetic scenario generation.

Evaluation Criteria

Auditability, reproducibility, false-positive usefulness, reviewer effort, privacy, and dataset balance.

Expected Deliverables

Governance proposal, provenance design, review checklist, calibration method, and example reviewed change.

Acceptance Criteria

  • Current assignment paths are documented.
  • Two governance models are compared.
  • Required evidence for changes is defined.
  • Label changes and score changes are treated separately.

Follow-Up Opportunities

May lead to provenance fields and locked evaluation manifests.

Cross-Repository Impact

Testkit and research directly; adapter and extension consume the result.

Complexity

Spike

Impact

High — safeguards product measurements.

Suggested Labels

spike, area: scoring, area: fixtures, area: evaluation

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions