The Quanfire blog

We build AI for professional-services firms — including our own domain models where it counts. Here's how we think, what we ship, and the evidence behind it.

Research & Engineering

How we build our own models, and how we measure them.

Research & Engineering

The retrieval number we took back before we published it

We built a cross-lingual retriever for EU law: give it a provision in one language and it finds the match in another, across German, English, Spanish, French and Italian. Our first run scored +126.3% on a split that leaked; we retracted it before publishing, and it is not a larger version of the number below, because it never measured generalisation at all. On a split that holds out whole units, the clean number is recall@1 0.2919 to 0.6480, a +122.0% gain with intervals that do not overlap. This time the check ran before the post did.

Aug 23, 2026
Research & Engineering

We trained the same model five more epochs. Then a clean split took the number back.

We published a 53% Recall@1 gain for this model. It was measured on a split that leaked, and we have withdrawn it. A clean re-run of the same corpus and recipe scores 0.1836 to 0.2518, +37.1%, intervals disjoint. Those are different weights: the ones this post announced were never re-measured and have no clean number. The per-language result changed too.

Aug 17, 2026
Research & Engineering

Harder training examples made our reranker worse. Here is the curve that fixed it.

We open-sourced rerank-gov-indic, a cross-lingual reranker for Indian government text. Correction (2026-08-23): the negative-hardness curve in this post was trained unseeded, so its "sweet spot" ordering is not established. See the correction note on the post.

Aug 16, 2026
Research & Engineering

We withdrew this reranker's +63.8%: it was measured on a split that leaked. The retrained model gains +47.2%.

We published a +63.8% Recall@1 gain for our first reranker. It was measured on a split that leaked, and we have withdrawn it. A retrained model, published as revision v2.0.0, was measured on a document-isolated split with 18 of 856 Central Acts held out entirely, and gains +47.2% (Recall@1 0.0723 to 0.1064, paired interval excludes zero). Those are different weights: the ones this post announced, v1.0.0, carry no valid number.

Aug 15, 2026
Research & Engineering

We said our legal model was weak on statutes. Here is the one that is not.

We open-sourced embed-statute-en, a 2.4 MB adapter for retrieving Indian central statutes. On the un-gameable low-overlap slice it more than doubles Recall@1 (+131%), with the numbers, confidence intervals, and provenance to check it yourself.

Aug 14, 2026
Research & Engineering

Retrieval across 16 Indian languages, and the receipts to check it

We open-sourced embed-gov-indic: a 2.4 MB adapter that retrieves Indian government press releases across 16 languages, +27.9% Recall@1 over the base model. The model is public and the method is documented, so you can pull it and judge it on your own text.

Aug 12, 2026
Research & Engineering

Don't take our word for it: read the tokens yourself

Quanfire's multilingual embeddings are live on Hugging Face: Indic-first, openly licensed, trained only on clean, documented data, verifiable in your browser.

Aug 11, 2026
Research & Engineering

The judgment is public domain. The headnote is not.

A court judgment is public domain; the headnote printed above it is not. How we excised the copyrighted layer so a legal embedding ships Apache-2.0.

Aug 10, 2026

For Practitioners

Where AI actually helps professional-services teams — and where it doesn't.