Internal Technical Memo

repo-state: Current Findings

What the prototype currently demonstrates, where retrieval fails, and what the next version should improve.

Prototype: repo-state v0.6.x Benchmark repo: ShitList Model: Jev latest Benchmark cases: 12 Questions: 36

repo-state is increasingly best understood as a deterministic repository indexing and retrieval layer, not as a replacement for Jev's judgment.

Current thesis: repo-state indexes; Jev decides.

1. Architecture

The useful comparison is not “Can repo-state reason better than Jev?” It is: Can Jev reach the same decision with a much smaller, query-specific representation of the repository?

CODEBASE
  ↓
deterministic scan
  ↓
compact deterministic index
  ↓
query-time retrieval
  ↓
hydrate relevant evidence
  ↓
Jev
  ↓
decision

Generative interpretation remains optional enrichment. It is not required for core indexing or retrieval.

2. Current Benchmark Result

Raw Jev accuracy
100%
Retrieved accuracy
94.44%
Average raw payload
86.5 KB
Average retrieved payload
15.8 KB
Average raw Jev latency
347 ms
Average retrieved Jev latency
153 ms
Average retrieval latency
2.1 ms
Index build latency
5.3 ms
Metric Raw repo Retrieved Interpretation
Accuracy 100% 94.44% Retrieval is not lossless yet.
Brier score 0.0117 0.0404 Calibration worsens when key evidence is omitted.
Decision payload 86.5 KB 15.8 KB Approximately 81.8% less context.
Jev latency 347 ms 153 ms Approximately 56% faster.
Retrieval latency — 2.1 ms Retrieval overhead is negligible.
Index build — 5.3 ms Cheap enough to refresh frequently.

3. What repo-state Is Responsible For

Jev remains responsible for questions such as whether a change weakens authentication, alters persistence, introduces risk, or requires review.

4. Where Retrieval Failed

Adding a task status

Raw Jev correctly determined that adding a new CANCELLED status required a schema change. The retrieved path did not.

Raw:       requires_schema_change = 0.75  ✓
Retrieved: requires_schema_change = 0.33  ✗

Retrieval selected public/app.js and server.js, but did not select the persistence/schema definition.

Adding a task priority

Raw:       requires_database_change = 0.55  ✓
Retrieved: requires_database_change = 0.15  ✗

Again, retrieval selected task and UI-related files but missed the database-level constraint.

Interpretation: the problem is not that 15 KB is inherently too little context. The retriever is choosing the wrong 15 KB for persistence-related questions.

5. Retrieval Quality Is Now the Primary Bottleneck

The current retriever is still too lexical. Terms such as task, status, and priority naturally rank application files highly, while the database file may contain the actual invariant Jev needs.

The next version should introduce typed relevance:

retrieval_score =
    lexical_match
  + file_role_match
  + architectural_type_match
  + relationship_proximity
  + direct_symbol_match

Example role mappings:

{
  "schema": [
    "schema",
    "database",
    "migration",
    "constraint",
    "column",
    "table"
  ],
  "persistence": [
    "database",
    "persist",
    "storage",
    "record"
  ],
  "auth": [
    "auth",
    "authentication",
    "authorization",
    "session",
    "cookie"
  ]
}

6. Index Pollution

Some queries retrieved generated or internal artifacts that should not normally compete with application source.

.repo-state/**
.meridian/transactions/**
benchmarks/**
coverage/**
node_modules/**
.git/**

These should be excluded from the default retrieval pool unless explicitly configured otherwise.

7. Claims We Can and Cannot Make

Claim Status
repo-state improves Jev accuracy. Not demonstrated.
repo-state improves Jev calibration overall. Not demonstrated.
repo-state can preserve most Jev decisions with substantially less context. Demonstrated.
repo-state materially reduces Jev decision latency. Demonstrated in this benchmark.
Deterministic indexing is cheap enough to refresh frequently. Demonstrated.
Retrieval quality is currently the main bottleneck. Strongly supported by observed failures.

8. Target for the Next Benchmark

RAW        100% accuracy   ~86 KB   ~347 ms
RETRIEVED  100% accuracy   ~16 KB   ~153 ms

If typed retrieval restores the two missed persistence decisions without materially increasing payload size, the clean claim becomes:

repo-state reduced the repository context Jev needed by approximately 82% and cut decision latency by approximately 56%, with no loss in decision accuracy.

9. Current Product Framing

A deterministic semantic index for codebases so Jev does not have to reason over the entire repository every time.

The longer-term value is not “AI understands your repo.” The value is that a decision system receives a compact, query-specific, evidence-backed view of the repo.