Composable audience intelligence · privacy-preserving data infrastructure
The data stays. The answer travels.
Every institution sits on data too dangerous to use — fraud signals, voices, customer records. Clean Lava is a composable audience intelligence platform: privacy-preserving infrastructure that extracts the answers and leaves the data where it was born — on the device, inside the institution, out of reach. On-device processing, data clean rooms, federated learning, differential privacy — decades-old mathematics, finally shipped as working rails. Built in Lagos, under NDPA 2023 and GDPR, for markets where getting this wrong isn't theoretical.
The pragmatic approach
No moonshots.
Working rails.
Privacy tech has a credibility problem: too many whitepapers, too few systems in production. We took the opposite bet.
Products before platform.
We don't sell a vision deck. The behavioral scoring engine is live — you can type into it right now and watch it recognize you. The platform will be earned, product by product, not promised.
Regulation is the design brief.
NDPA 2023 and GDPR aren't compliance chores bolted on before launch. They are the specification we build against — which is why what we ship passes legal review instead of negotiating with it.
Boring math, radical outcomes.
Differential privacy, private set intersection, federated learning — peer-reviewed, decades old, gloriously unglamorous. We compose proven primitives carefully instead of inventing new risk.
Built where the stakes are highest.
Fraud in Lagos doesn't forgive sloppy engineering, and Nigerian regulators don't grade on a curve anymore. Infrastructure hardened here travels anywhere.
The privacy stack
Five primitives.
One promise: raw data doesn't move.
01 / ON-DEVICE PROCESSING
The computation commutes. The data doesn't.
Instead of hauling behavioral signals, keystrokes, and voice into a warehouse to analyze them, we send the analysis to the device that produced them. Raw events live and die in local memory; only a handful of summary numbers ever touch a network.
Unlocks: fraud verdicts and language understanding from data that is never collected in the first place.
02 / DIFFERENTIAL PRIVACY
Answers that can't betray anyone.
Every score and every aggregate we release carries calibrated mathematical noise — enough to make reconstructing any individual statistically implausible, small enough to leave the answer useful. It's a provable property of the output, not a promise in a policy.
Unlocks: publishable, shareable intelligence from populations of one million or one.
03 / PRIVATE SET INTERSECTION
Two datasets meet in the dark.
PSI lets two organizations discover what their encrypted datasets have in common — shared fraud cases, shared customers — without either side decrypting, seeing, or receiving the other's records. The overlap is computed; the lists never exist in one place.
Unlocks: "how much do we overlap?" answered between competitors, legally.
04 / FEDERATED LEARNING
The model travels. The data stays home.
To improve a model, we send it to where the data lives — a bank's infrastructure, a fleet of handsets — let it learn locally, and bring back only the weight updates. A fraud model can learn from ten banks' data while no bank's data ever leaves its walls.
Unlocks: models trained on data that could never be pooled — because it never is.
05 / THE DATA CLEAN ROOM
A room where records can't leave.
The primitives compose into a neutral, audited environment where multiple parties bring encrypted data, computation happens blind, and the only thing that exits is an aggregate answer — noised, minimum-cohort-sized, and useless for re-identification. We operate the room; we can't see into it either.
Unlocks: Share Without Sharing — collaboration your legal team can sign.
Product lines
Three ways to use data
you'd never dare hand over.
Each product is a question institutions currently can't ask without collecting something they shouldn't hold. Same stack underneath; different doors it opens.
Behavioral Security
Any attacker can steal a password. None of them can steal the way your customer types. Lava Sense scores every session against the account owner's behavioral signature — on-device, in milliseconds — and hands your server one word: ALLOW, STEP_UP, or BLOCK.
See it work, liveLanguage Intelligence
Two thousand African languages; almost none of them modeled. We build on-device voice and text understanding from consented, revocable, native-speaker corpora — and improve it by federated learning, so the model gets smarter while the voices never leave the device.
The possibilitiesShare Without Sharing
Two institutions. One question — shared fraud, shared customers, campaign truth. Our clean room computes the answer on encrypted inputs via private set intersection, and the answer is the only thing that walks out.
The possibilitiesAnatomy of a verdict
From keystroke to verdict
in three moves.
Capture at the source
The SDK measures how a person types and moves — never what they type. Raw events exist only in the browser's memory.
Reduce before it moves
Behavior collapses to a compact set of timing statistics on the device — session summaries plus per-key averages under opaque labels. The raw buffers are discarded in the same breath.
Score, noise, answer
The API scores the session against the device's enrolled baseline, applies differential-privacy noise, and returns a verdict your risk policy can act on.
Built for your world
Find your situation.
See what becomes possible.
The same primitives unlock different doors in different industries. Pick yours.
Account takeover is a behavior problem. Treat it like one.
SIM-swap and credential-stuffing attacks pass every password check you have, because the attacker has the credentials. What they can't fake is the customer's hands.
- Score every login and transfer for takeover — continuous behavioral authentication without storing a single biometric record.
- Step up only the sessions that deserve it — three-state verdicts mean OTP friction lands on attackers, not customers.
- Compare fraud lists with a rival bank — in a clean room, via private set intersection: overlap out, records never.
Delivered by → Behavioral Security · Share Without Sharing
Ship devices that understand your customers — in their own language.
Global voice assistants don't speak Yoruba, Ijebu, or Sheng. And network-level fraud — SIM-box, churn-and-burn — hides in signals you hold but can't legally pool.
- License on-device language models — voice and text understanding for African languages, running on the handset, no cloud round-trip.
- Detect SIM-box fraud across networks — pool fraud signals with other operators in a clean room; get the pattern, never the subscriber data.
- Anchor behavioral security at the hardware layer — OEM-level typing and handling signatures that apps can build on.
Delivered by → Language Intelligence · Share Without Sharing
Underwrite people the credit bureau has never met.
Most of your addressable market has no formal credit history. The behavioral signal exists — but collecting it raw would be a regulatory and ethical liability.
- Behavioral risk scores from consented usage patterns — underwriting signal without warehousing anyone's personal data.
- Catch application fraud at the form — bots and coached applicants type differently than genuine ones.
- Measure portfolio overlap with a reinsurer — aggregate answers from a clean room, no policyholder lists exchanged.
Delivered by → Behavioral Security · Share Without Sharing
The cookie is gone. The signal doesn't have to be.
Third-party identifiers are dead and African ad budgets are flying blind — no reliable audience signal, no clean attribution, and bots inflating every metric.
- Privacy-safe audience segments — behavioral cohorts built on-device, released only as differentially-private aggregates.
- Attribution without identity leaks — match conversions in a clean room aligned with IAB PAIR v1.1.
- Cut bot traffic out of your funnel — behavioral scoring separates automation from humans before it poisons your data.
Delivered by → Behavioral Security · Share Without Sharing
Fraud crosses platforms. Your defenses should too.
Mule networks rotate across remittance corridors faster than any single operator can blacklist them — and no operator can legally hand its customer list to a competitor.
- Cross-platform fraud clean rooms — find the accounts hitting three corridors at once, without sharing who your customers are.
- Behavioral checks on high-value transfers — a takeover of a diaspora account looks wrong at the keyboard before it looks wrong on the ledger.
- Diaspora credit signals — consented, aggregate-only indicators for cross-border lending.
Delivered by → Share Without Sharing · Behavioral Security
Population intelligence without a surveillance state.
Agencies need cross-ministry answers — health, migration, subsidy fraud — but centralizing citizen records creates exactly the honeypot the NDPA exists to prevent.
- Cross-agency statistics, not cross-agency databases — each ministry keeps its records; the clean room computes only the aggregate answer.
- Subsidy-fraud detection with due process — flag anomalous claim patterns without profiling individual citizens.
- Language access for public services — on-device models that let citizens interact with services in their own language.
Delivered by → Share Without Sharing · Language Intelligence
Applications & possibilities
One infrastructure,
many markets.
Everything below runs on the same rails. Outcome first, mechanism second.
Account-takeover prevention
Stop SIM-swap and credential-stuffing at the session, via continuous on-device behavioral scoring.
Bot mitigation
Separate humans from automation in signups, checkouts, and ad funnels — behavior is harder to spoof than headers.
Cross-organization fraud lists
Pool fraud intelligence across institutions through private set intersection; overlap counts out, records never.
Behavioral underwriting
Risk scores for thin-file customers from consented usage patterns — crop insurance, PAYG solar, informal commerce.
Federated analytics & attribution
Audience measurement and campaign attribution computed where the data lives, released as differentially-private aggregates.
African-language interfaces
Voice and text understanding for local languages and dialects, running entirely on the device that heard them.
Why trust it
Privacy isn't a policy here.
It's the physics.
Every design decision starts from one question: how do we make this impossible to abuse — including by us?
- On-device processing. Raw keystrokes, pointer paths, and voice never leave the device. Reduction to summary features happens locally, and the buffers are discarded.
- Differential privacy on every output. Calibrated noise is added to every score before it leaves our infrastructure, so no individual can be reverse-engineered from any answer.
- No raw data movement — architecturally. Our API has no endpoint that accepts raw behavioral events for storage. The guarantee is structural, not contractual.
- Regulated from day one. NDPA 2023 and GAID 2025 at home, GDPR for global partners: explicit consent, data minimization, right to erasure — including a one-call baseline reset.
Common questions
Asked by every serious buyer.
Do you collect personally identifiable information?
No. The SDK never reads what a user types — only timing and motion patterns, which are reduced in the browser to a small set of timing statistics under opaque, device-salted labels. Device identity is a locally-generated random ID, hashed with SHA-256 before it leaves the device. There is no fingerprinting, no IMEI, no contact with the user's identity.
What happens if a legitimate customer gets flagged?
The verdict model is three-state by design: most uncertainty resolves to STEP_UP (ask for one more factor) rather than BLOCK. Low-risk sessions continuously refine the baseline, so a customer whose behavior drifts — new keyboard, injured hand — re-converges instead of being locked out.
How is this different from device fingerprinting?
Fingerprinting identifies a device across sites, usually without consent — it's a tracking technology. Lava Sense verifies that the person behind an already-authenticated session behaves like the account owner, inside one application, with consent, and with differential-privacy noise on every score. It answers "is this them?", never "who is this?".
Design partners
The next piece of this
gets built with someone. You?
We're onboarding a small group of design partners — fraud, risk, data, and partnerships leaders at banks, telcos, insurers, and platforms.
Prefer email? Partnerships: partnerships@cleanlava.com