← All posts
Research & Engineering

We withdrew this reranker's +63.8%: it was measured on a split that leaked. The retrained model gains +47.2%.

We withdrew this reranker's +63.8%: it was measured on a split that leaked. The retrained model gains +47.2%.
A correction, added August 23, 2026. We are withdrawing the entire measurement in this post, not one figure from it. The two sections below, "The number that matters" and "The headroom we are not hiding", were produced from a single evaluation split that excluded nothing: every one of the 1,200 held-out queries shared a document with the training data, and most of the held-out answers appeared in that training data word for word. Everything those sections report came from that split: the 0.205 and 0.336 in the table, the +0.131 gain, the +63.8%, the paired confidence interval, the claim that it excludes zero, and the 0.718 retrieval ceiling. A measurement made on a split that excluded nothing measures nothing, so none of those numbers stands, and the difference between the two table rows is not a result either. There is nothing here to revise down. We are taking the measurement out. Since then we retrained this reranker on the clean data and measured the new weights on a split with the document isolation the first evaluation lacked: 1,494 queries over a 1,205-passage pool, with 18 of 856 Central Acts held out entirely, so that no Act the model trained on appears anywhere in the evaluation. That holdout by whole Act is the reason this number means what it says. On that split, retrieve-then-rerank moves Recall@1 from 0.0723 [0.0589, 0.0857] to 0.1064 [0.0910, 0.1218], a gain of +0.0341, or +47.2%, with a paired bootstrap 95% confidence interval of [+0.0207, +0.0482] that excludes zero. The retrieval ceiling on this split is 0.7557, and of the queries whose gold passage is recoverable in the retriever's top-100, the reranker lifts 14.1% into first place. These are lower absolute numbers than the withdrawn section showed, because this is a harder and more honest test, not a worse model. It is a different evaluation set reported from zero. It is not a revised version of the withdrawn figures, the two are not comparable, and we will not present them as a before-and-after. That +47.2% belongs to a new set of weights, published as revision v2.0.0. The weights this post originally announced are revision v1.0.0, and they are the ones the withdrawn measurement was taken on. They carry no valid effectiveness number, not a lower one, none. If you pulled this model when the post first went up and pinned the revision, you are holding v1.0.0. Move to v2.0.0 to get the model the +47.2% describes. One limit belongs on that new number. It comes from a single training run, and that run is unseeded: the split, the validation carve and the confidence interval are seeded and repeatable, but the training itself is not, so a re-run would begin from a different initialisation and a different draw of training examples and we cannot promise it lands on the same value. What we are stating is the one measurement we made. What we are not claiming is that the improvement is a stable property of the procedure, because we have not measured its run-to-run spread. The +47.2% is the honest measurement of these weights; it is not withdrawn by that limit. What survives is narrower than the post below claims and still real: rerank-statute-en at v2.0.0 is a genuine, open, Apache-2.0 cross-encoder that reorders a retriever's shortlist on English central-statutory text, and on a split with no leakage it improves end-to-end Recall@1 over that retriever by +47.2%. The rest of this post was written around the withdrawn measurement and reads as if that measurement were established. Treat everything below this notice as the original post, kept for the record, and this paragraph as what we now stand behind.

Retrieval is only half of search. A retriever casts a wide net and pulls back a shortlist of candidate passages. Something still has to read that shortlist closely and decide which one actually answers the query. That second step is reranking, and today we are open-sourcing our first one.

It is called rerank-statute-en, it is public on Hugging Face, and the weights are Apache-2.0.

huggingface.co/quanfire-ai/rerank-statute-en

What it is

rerank-statute-en is a cross-encoder reranker for English central-statutory, bare-Act text. Where a retriever encodes the query and each passage separately and compares vectors, a cross-encoder reads the query and one candidate passage together and returns a single relevance score. You use it as a second stage: a bi-encoder retriever pulls the top-k candidates, and the reranker reorders them so the best answer moves to the top.

It is the companion to embed-statute-en, the statute retriever we released earlier. That model finds candidates; this one decides the order. It is a reranker, not a retriever, and it cannot stand in for one.

The number that matters

We measured the honest end-to-end quantity: retrieve first, then rerank, and see how often the single top result is correct. On 1,200 held-out queries against 1,182 unique statutory passages, scored on GPU:

StageRecall@195% CI
Retrieve only (bi-encoder)0.205[0.182, 0.229]
Retrieve then rerank0.336[0.310, 0.363]

That is a +0.131 gain, or +63.8%, and the paired 95% confidence interval is [+0.106, +0.157], which excludes zero. The improvement is not noise: on the same queries, adding the reranker moves the right passage into first place far more often than not.

The headroom we are not hiding

An absolute Recall@1 of 0.336 is modest, and we will say so plainly. The reason is a ceiling we do not control: the retriever only surfaces the correct passage anywhere in its top-100 about 71.8% of the time. A reranker can only reorder what it is handed, so it can never rank a passage the retriever never retrieved. That 0.718 is the wall this stage is working under. Getting from 0.205 to 0.336 against it is a real, significant first step, and there is genuine headroom left. This is a strong first reranker, not a claim that statute search is solved.

The bug worth telling you about

The first training run collapsed. The reranker scored candidates almost at random, no better than no reranker at all. The cause was in the negatives, not the model.

To train a reranker you show it correct query-passage pairs and a pile of wrong ones, and it learns to score the right pair higher. We first drew those wrong pairs with ordinary hard-negative mining, and it turned out the negatives were form-separable: they carried surface tells, things like OCR artifacts and repeated boilerplate, that let the model tell right from wrong without reading the query at all. It learned the shortcut instead of the task, and a shortcut does not transfer to real candidates at inference time.

The fix was to change the negatives, not the architecture. We switched to form-matched negatives: clean statutory bodies drawn from other records, so a wrong candidate looks exactly like a right one except that it does not answer the query. With the shortcut gone, the only way to score well is to actually compare meaning, and the model learned to do that. The lesson generalizes: a reranker is only as honest as its negatives, and those negatives have to match the candidate distribution it will see at inference. If your wrong answers are distinguishable by anything other than relevance, your reranker will learn that anything else.

Where the data came from

The reranker is trained on 858 Central Acts of the Indian Parliament, from a public dataset (Zenodo record 5088102, released under CC-BY-4.0). We attribute it plainly:

Source: "An annotated dataset of Central Acts enacted by the Indian Parliament," Zenodo record 5088102, CC-BY-4.0.

The model is non-reconstructive: it emits a relevance score and cannot reproduce the source text. We ship the weights, not the corpus. Reproduction of a bare Act is a permitted act under section 52(1)(q)(ii) of the Copyright Act, 1957. This is general information about how the model was built, not legal advice.

What it is for, and what it is not

We are precise about scope. rerank-statute-en is built and validated for English central-statutory (bare-Act) text, as a second-stage reranker over a retriever's shortlist. It is not a retriever and cannot be used as one. It is not validated for court judgments, for state legislation, rules, regulations or notifications, or for non-English text. A model that tells you where it works is worth more than one that claims to work everywhere.

Why we build in the open

This is another clean, verifiable model put out in the open, and the reason we told you about the collapsed run is the same reason we publish the confidence intervals: we would rather you trust the method than the marketing. The model card has the full benchmark, the ceiling, the intervals and the scope in one place. Pull it, put it behind your own retriever, and judge it on your own statutory text. Our training framework is open source at github.com/quanfire-ai/quanfire-multilingual-embedding (Apache-2.0). That is the discipline we bring to our products too.

DocPro

Want that same measure-it-honestly discipline applied to your own documents? The same team builds DocPro. Try it on your contracts and filings.

Share

Keep reading