EVIDENCE
Benchmarks
Geltre claims are scoped to measured artifacts. Positive results and failed gates are both part of the public technical story.
ServiceOps v0.3
| Selector | Required-evidence recall @ 5 | JEV accuracy |
|---|---|---|
| BM25 | 0.6255 | 0.8235 |
| Geltre linear | 0.9961 | 1.0000 |
| Geltre neural | 0.9647 | 1.0000 |
The linear ServiceOps ranker is the current reference because it matched the downstream neural result while being smaller and faster.
CodeOps cross-distribution v0.8
Negative result: the predeclared robustness gate did not pass.
| Metric | Actual | Required |
|---|---|---|
| Required-evidence recall | 0.808219 | ≥ 0.98 |
| Context reduction | 0.927891 | ≥ 0.50 |
The consumed v0.8 holdout is development evidence only. A broader production claim requires a new untouched successor cohort.
Release principle
Geltre must reduce context, cost, or distraction while preserving downstream task success better than ordinary retrieval. Context reduction alone is not sufficient.