§Research
Three questions,
asked carefully.
Our research grows out of systems we run in production. We publish methods, not live results, and nothing leaves the lab until it has been measured.
Area 01
Forecasting & calibration
When a system says 70%, does it happen 70% of the time?
Probabilistic forecasts are only useful if their numbers mean what they say. We work on calibration, making 70% mean 70%, and on evaluation that stays honest when the data is messy, duplicated and adversarial.
Questions we’re working on
- How far does a forecast’s confidence overstate what it knows, and where does it overstate most?
- How should forecasts be scored when the same event is predicted many times over?
- What evidence should a system need before it trusts an apparent advantage?
Methods
- Calibration
- Proper scoring rules
- Closing-line value
- Anytime-valid testing
- Out-of-sample gates
Area 02
Market measurement
How well do real markets price risk?
Prices are forecasts too. We measure how the markets people actually use set and correct their prices: the margin they charge, how far books disagree, and whether a single book’s prices agree with each other. We start with sportsbooks serving African bettors, a region the market-efficiency literature has largely skipped.
Questions we’re working on
- How much margin do these markets charge, and where is it hidden?
- When books disagree, is the gap large enough to matter?
- When a price looks wrong, is the market wrong, or the data?
Methods
- Margin removal
- Cross-book dispersion
- Coherence checks
- Large-scale price capture
Area 03
Trustworthy AI systems
Can a language model be trusted with facts that matter?
Language models are fluent, and fluency is not accuracy. We design systems where the model reads and reasons, deterministic code owns every fact and action that matters, and a separate checker verifies claims against the evidence before anyone sees them.
Questions we’re working on
- How do you stop an agent from inventing a number?
- When should a verifier repair a claim, and when should it refuse?
- How do you turn messy, code-switched conversation into safe, typed actions?
Methods
- Agentic evidence gathering
- Verified extraction
- Calibrated judges
- Typed commands
§Method
How we decide what’s true.
The same rules apply to a paper, a product feature and a page on this site. They were learned the hard way, on our own work first.
Box 1 · The discipline
Measured, or it doesn’t exist.
Every claim names its data, its sample size and its method.
Count what’s independent.
Statistics are computed on independent events, never inflated by repeated rows of the same one.
One method per claim.
A result that only appears under a favourable method is a bug report about the method.
Gates are allowed to say no.
An evaluation that admits nothing isn’t broken. Empty can be the honest answer.
Abstain rather than guess.
What we can’t resolve, we name. We don’t estimate it.
§Publications
Publications
Posted to arXiv first, then submitted to journals and conferences. Nothing is listed here before it has a timestamp.
| Title | Area | Venue | Year | Status |
|---|---|---|---|---|
No entries yet.The first manuscripts are in preparation. | ||||
§Correspondence
Working on similar questions?
We read every message about the research, including the ones that tell us we’re wrong.
Tafiti Labs · Lagos, Nigeria