We trained the same model five more epochs. Then a clean split took the number back.

A correction, added 08/23/2026. After we published this post we re-ran the evaluation behind it on a document-isolated split, one where no passage from the training data can appear in the evaluation set. Being exact about what that re-run was matters more than the number it produced: it trained a second set of weights, and the clean number below belongs to those. They are published separately, as revision clean-2026-08-21. The weights this post announced, the ones on the default branch, have no clean measurement at all. On the clean split the re-run improves on the base model: cross-lingual Recall@1 goes from 0.1836 (95% CI [0.1651, 0.2038]) to 0.2518 ([0.2307, 0.2741]), a +37.1% relative gain, with intervals that do not overlap. Recall@5, Recall@10, MRR and nDCG@10 were re-measured on the same split and are published beside it on the model card. What every one of those numbers has in common is the thing this post cannot give you: they measure the re-run, not the weights this post announced. Two things follow, and the second is the larger. The first is the headline. The "+53%" in the title and throughout this post was measured on the split we no longer trust, and it described the released weights. It has no clean number of any size to replace it, because those weights were never re-run. The +37.1% above is not a corrected +53%: it belongs to the second set of weights, not to what this post announced, and we will not present the two as a before and after. The second reaches the argument of the post itself. The case we made here was that training longer helped, that v1.1 improves on our own v1.0. We never re-ran v1.0, or the earlier points, on the clean split, so that comparison has no clean measurement on either side. And the curve that carried the argument is mislabelled. Its four points are separate training runs, not epochs of one run. Three of the four were trained for a single epoch each and only the last for six; the earlier runs also differ from one another in their corpus and in the kind of pairs they were built from, not only in how long they ran. So "the improvement from longer training," drawn as one variable rising smoothly, was never the experiment the curve depicts. This holds whether or not the split was clean. The per-language picture also changed, and not only in size. On the clean split, Odia improves (0.138 to 0.181) where this post reported it regressing, and Manipuri falls to 0.000. Measured against the base model, three languages come out below it: Manipuri, Gujarati and Khasi, the last on only 14 evaluation queries. Twelve of the sixteen improve and Nepali, on seven queries, is flat. What this post claimed was that thirteen improved and Gujarati and Odia regressed, measured against v1.0 on the split we no longer trust. The clean per-language numbers are on the model card. The "top ten about 84% of the time" figure was re-measured too, and it also moved: on the clean split Recall@10 is 0.7521, about 75 percent, against the 0.843 this post reports. Like the headline, it belongs to the re-run and not to the weights this post announced. What survives cleanly is narrower and still real: this is a genuine, open, Apache-2.0 cross-lingual retriever for Indian government press releases, and on a split with no leakage the same corpus and recipe beat the base model by +37.1%. What this post can no longer tell you is by how much the weights we actually released beat it, because we never re-measured them. The claim that more training beat less training is one we have to earn again before we make it.Our first release of this model was undertrained. We shipped it at a single epoch, measured it honestly, and it was a real improvement over the base model. Then we let the same model keep training on the same data, measured every epoch along the way, and it kept getting better. This is the story of what five more epochs bought us, and the honest part we are not hiding: two of the sixteen languages came out worse.
The model is embed-gov-indic v1.1, revision gov-indic-e4-ep6. It is public on Hugging Face and the weights are Apache-2.0.
What it is
embed-gov-indic is a cross-lingual retriever for Indian government press-release text across 16 Indian languages. Cross-lingual means you can ask in one Indian language and it will find the release about the same event that was filed in another. It is built on the open multilingual-e5-small base model, and it encodes a query and each passage into vectors so the closest passage to your query rises to the top. You give it a search query, it gives you back the releases most likely to answer it.
It is a specialist. We built and measured it for Indian government press releases, and that is the only thing we claim it does well. It is not a general-purpose search model and it is not a legal model. A model that tells you where it works is worth more than one that says it works everywhere.
The number that matters
We measure the plain, useful quantity: how often the single top-ranked result is the correct one. That is Recall@1. We score it on an in-distribution, cross-lingual query set, on GPU, and we compare three checkpoints on the exact same evaluation.
| Model | Cross-lingual Recall@1 |
|---|---|
| Base (multilingual-e5-small) | 0.184 |
| embed-gov-indic v1.0 (one epoch) | 0.235 |
| embed-gov-indic v1.1 (six epochs) | 0.281 |
From the base model to v1.1 is 0.184 to 0.281, a 53% relative gain, and the bootstrap 95% confidence intervals for the two do not overlap. The same holds when we compare v1.1 against our own v1.0: the improvement from longer training is real and its interval clears zero. Put the other way, the correct passage lands in the top ten about 84% of the time.
One caution we will repeat wherever these numbers appear: Recall@1 and top-ten are retrieval ranking measures. They describe how often the right document is placed first, or near the top. They are not an accuracy or a confidence percentage, and we do not translate them into one.
The curve, epoch by epoch
The reason we can tell you the first version was undertrained is that we scored every epoch, not just the one we shipped. Here is the whole arc, as relative gain over the base model:
- Epoch 1: +8.7%, confidence interval overlaps the base, so we treat it as not yet distinguishable.
- Epoch 2: +16.2%, still overlapping.
- Epoch 3, released as v1.0: +27.7%, and here the interval first clears the base.
- Epoch 4, released as v1.1: +53%, clearing both the base and v1.0.
The lift did not arrive all at once, and the early epochs did not yet separate from the base with any confidence. We shipped v1.0 at the first point where the gain was solid, and we shipped v1.1 when more training made it solidly better again. If we had only looked at the final model, we could not have told you that.
The honest part: two languages got worse
More training helped in aggregate, but not everywhere. Going from v1.0 to v1.1, thirteen of the sixteen languages improved, and two regressed: Gujarati and Odia. They are better under v1.0 than under v1.1.
We could have replaced v1.0 and said nothing. Instead we kept both, each under its own revision tag, so anyone who cares about Gujarati or Odia can pin whichever version serves them better. Telling you which two languages moved the wrong way is the point, not a footnote to bury.
Two more scope limits belong here in the open. The headline is a single aggregate across a mixed-language query set, not a per-language guarantee. And our lowest-resource languages, Khasi, Nepali and Manipuri, are thin in the training data, so treat them as the least reliable of the sixteen.
Where the data came from
embed-gov-indic is trained on Press Information Bureau (PIB) government press releases. The same release is often published in many languages, which is exactly what gives us naturally parallel cross-lingual pairs to train on. We reuse that text under PIB's reproduction policy, which permits reproduction with attribution and carries no non-commercial and no share-alike restriction. The training is non-reconstructive: the model learns to place text in a vector space and cannot reproduce the source releases. We ship the weights, not the corpus, and the weights are clean Apache-2.0. This is general information about how the model was built, not legal advice.
Verify it yourself
We would rather you trust the method than the marketing, so everything above is checkable. The Hugging Face model card carries the full benchmark, the per-version numbers and the scope in one place. And the retriever is wired into our live cross-lingual playground demo, so you can type a query in one Indian language and watch it surface a government release filed in another. Pull the model, run it against your own text, and judge it on what you see.
huggingface.co/quanfire-ai/embed-gov-indic
Want that same measure-it-honestly discipline applied to your own documents? The same team builds DocPro. Try it on your contracts and filings.