repo-state: Current Findings
What the prototype currently demonstrates, where retrieval fails, and what the next version should improve.
repo-state is increasingly best understood as a deterministic repository indexing and retrieval layer, not as a replacement for Jev's judgment.
1. Architecture
The useful comparison is not “Can repo-state reason better than Jev?” It is: Can Jev reach the same decision with a much smaller, query-specific representation of the repository?
CODEBASE ↓ deterministic scan ↓ compact deterministic index ↓ query-time retrieval ↓ hydrate relevant evidence ↓ Jev ↓ decision
Generative interpretation remains optional enrichment. It is not required for core indexing or retrieval.
2. Current Benchmark Result
| Metric | Raw repo | Retrieved | Interpretation |
|---|---|---|---|
| Accuracy | 100% | 94.44% | Retrieval is not lossless yet. |
| Brier score | 0.0117 | 0.0404 | Calibration worsens when key evidence is omitted. |
| Decision payload | 86.5 KB | 15.8 KB | Approximately 81.8% less context. |
| Jev latency | 347 ms | 153 ms | Approximately 56% faster. |
| Retrieval latency | — | 2.1 ms | Retrieval overhead is negligible. |
| Index build | — | 5.3 ms | Cheap enough to refresh frequently. |
3. What repo-state Is Responsible For
- Map files, routes, dependencies, symbols, schemas, and relationships.
- Build a compact reusable repository index.
- Given an intent and a question, identify the most relevant parts of the repo.
- Hydrate only the evidence needed for the current decision.
- Remain deterministic by default.
Jev remains responsible for questions such as whether a change weakens authentication, alters persistence, introduces risk, or requires review.
4. Where Retrieval Failed
Adding a task status
Raw Jev correctly determined that adding a new CANCELLED status required a schema change.
The retrieved path did not.
Raw: requires_schema_change = 0.75 ✓ Retrieved: requires_schema_change = 0.33 ✗
Retrieval selected public/app.js and server.js, but did not select the persistence/schema definition.
Adding a task priority
Raw: requires_database_change = 0.55 ✓ Retrieved: requires_database_change = 0.15 ✗
Again, retrieval selected task and UI-related files but missed the database-level constraint.
5. Retrieval Quality Is Now the Primary Bottleneck
The current retriever is still too lexical. Terms such as task, status, and
priority naturally rank application files highly, while the database file may contain the actual
invariant Jev needs.
The next version should introduce typed relevance:
retrieval_score =
lexical_match
+ file_role_match
+ architectural_type_match
+ relationship_proximity
+ direct_symbol_match
Example role mappings:
{
"schema": [
"schema",
"database",
"migration",
"constraint",
"column",
"table"
],
"persistence": [
"database",
"persist",
"storage",
"record"
],
"auth": [
"auth",
"authentication",
"authorization",
"session",
"cookie"
]
}
6. Index Pollution
Some queries retrieved generated or internal artifacts that should not normally compete with application source.
.repo-state/** .meridian/transactions/** benchmarks/** coverage/** node_modules/** .git/**
These should be excluded from the default retrieval pool unless explicitly configured otherwise.
7. Claims We Can and Cannot Make
| Claim | Status |
|---|---|
| repo-state improves Jev accuracy. | Not demonstrated. |
| repo-state improves Jev calibration overall. | Not demonstrated. |
| repo-state can preserve most Jev decisions with substantially less context. | Demonstrated. |
| repo-state materially reduces Jev decision latency. | Demonstrated in this benchmark. |
| Deterministic indexing is cheap enough to refresh frequently. | Demonstrated. |
| Retrieval quality is currently the main bottleneck. | Strongly supported by observed failures. |
8. Target for the Next Benchmark
RAW 100% accuracy ~86 KB ~347 ms RETRIEVED 100% accuracy ~16 KB ~153 ms
If typed retrieval restores the two missed persistence decisions without materially increasing payload size, the clean claim becomes:
9. Current Product Framing
A deterministic semantic index for codebases so Jev does not have to reason over the entire repository every time.
The longer-term value is not “AI understands your repo.” The value is that a decision system receives a compact, query-specific, evidence-backed view of the repo.