Composable audience intelligence · privacy-preserving data infrastructure

The data stays. The answer travels.

Every institution sits on data too dangerous to use — fraud signals, voices, customer records. Clean Lava is a composable audience intelligence platform: privacy-preserving infrastructure that extracts the answers and leaves the data where it was born — on the device, inside the institution, out of reach. On-device processing, data clean rooms, federated learning, differential privacy — decades-old mathematics, finally shipped as working rails. Built in Lagos, under NDPA 2023 and GDPR, for markets where getting this wrong isn't theoretical.

The pragmatic approach

No moonshots.
Working rails.

Privacy tech has a credibility problem: too many whitepapers, too few systems in production. We took the opposite bet.

Products before platform.

We don't sell a vision deck. The behavioral scoring engine is live — you can type into it right now and watch it recognize you. The platform will be earned, product by product, not promised.

Regulation is the design brief.

NDPA 2023 and GDPR aren't compliance chores bolted on before launch. They are the specification we build against — which is why what we ship passes legal review instead of negotiating with it.

Boring math, radical outcomes.

Differential privacy, private set intersection, federated learning — peer-reviewed, decades old, gloriously unglamorous. We compose proven primitives carefully instead of inventing new risk.

Built where the stakes are highest.

Fraud in Lagos doesn't forgive sloppy engineering, and Nigerian regulators don't grade on a curve anymore. Infrastructure hardened here travels anywhere.

The privacy stack

Five primitives.
One promise: raw data doesn't move.

01 / ON-DEVICE PROCESSING

The computation commutes. The data doesn't.

Instead of hauling behavioral signals, keystrokes, and voice into a warehouse to analyze them, we send the analysis to the device that produced them. Raw events live and die in local memory; only a handful of summary numbers ever touch a network.

Unlocks: fraud verdicts and language understanding from data that is never collected in the first place.

02 / DIFFERENTIAL PRIVACY

Answers that can't betray anyone.

Every score and every aggregate we release carries calibrated mathematical noise — enough to make reconstructing any individual statistically implausible, small enough to leave the answer useful. It's a provable property of the output, not a promise in a policy.

Unlocks: publishable, shareable intelligence from populations of one million or one.

03 / PRIVATE SET INTERSECTION

Two datasets meet in the dark.

PSI lets two organizations discover what their encrypted datasets have in common — shared fraud cases, shared customers — without either side decrypting, seeing, or receiving the other's records. The overlap is computed; the lists never exist in one place.

Unlocks: "how much do we overlap?" answered between competitors, legally.

04 / FEDERATED LEARNING

The model travels. The data stays home.

To improve a model, we send it to where the data lives — a bank's infrastructure, a fleet of handsets — let it learn locally, and bring back only the weight updates. A fraud model can learn from ten banks' data while no bank's data ever leaves its walls.

Unlocks: models trained on data that could never be pooled — because it never is.

05 / THE DATA CLEAN ROOM

A room where records can't leave.

The primitives compose into a neutral, audited environment where multiple parties bring encrypted data, computation happens blind, and the only thing that exits is an aggregate answer — noised, minimum-cohort-sized, and useless for re-identification. We operate the room; we can't see into it either.

Unlocks: Share Without Sharing — collaboration your legal team can sign.

Anatomy of a verdict

From keystroke to verdict
in three moves.

THE DEVICE — trust boundary Raw signals keystroke timing · pointer paths On-device reduction timing statistics · buffers discarded features only Clean Lava API baseline enrollment → live scoring differential-privacy noise on scores no raw data stored — ever verdict Your server ALLOW · STEP_UP · BLOCK gate the transaction, not the user

Capture at the source

The SDK measures how a person types and moves — never what they type. Raw events exist only in the browser's memory.

Reduce before it moves

Behavior collapses to a compact set of timing statistics on the device — session summaries plus per-key averages under opaque labels. The raw buffers are discarded in the same breath.

Score, noise, answer

The API scores the session against the device's enrolled baseline, applies differential-privacy noise, and returns a verdict your risk policy can act on.

Built for your world

Find your situation.
See what becomes possible.

The same primitives unlock different doors in different industries. Pick yours.

Banks & fintechs

Account takeover is a behavior problem. Treat it like one.

SIM-swap and credential-stuffing attacks pass every password check you have, because the attacker has the credentials. What they can't fake is the customer's hands.

  • Score every login and transfer for takeover — continuous behavioral authentication without storing a single biometric record.
  • Step up only the sessions that deserve it — three-state verdicts mean OTP friction lands on attackers, not customers.
  • Compare fraud lists with a rival bank — in a clean room, via private set intersection: overlap out, records never.

Delivered by → Behavioral Security · Share Without Sharing

0
raw behavioral records stored, ever
8.23%
equal-error rate on the public CMU keystroke benchmark — beats the best published detector
3
verdict states: ALLOW, STEP_UP, BLOCK
100%
of raw-signal processing happens on the device

Applications & possibilities

One infrastructure,
many markets.

Everything below runs on the same rails. Outcome first, mechanism second.

Fraud & identity

Account-takeover prevention

Stop SIM-swap and credential-stuffing at the session, via continuous on-device behavioral scoring.

Traffic quality

Bot mitigation

Separate humans from automation in signups, checkouts, and ad funnels — behavior is harder to spoof than headers.

Collaboration

Cross-organization fraud lists

Pool fraud intelligence across institutions through private set intersection; overlap counts out, records never.

Credit & risk

Behavioral underwriting

Risk scores for thin-file customers from consented usage patterns — crop insurance, PAYG solar, informal commerce.

Measurement

Federated analytics & attribution

Audience measurement and campaign attribution computed where the data lives, released as differentially-private aggregates.

Language

African-language interfaces

Voice and text understanding for local languages and dialects, running entirely on the device that heard them.

Why trust it

Privacy isn't a policy here.
It's the physics.

Every design decision starts from one question: how do we make this impossible to abuse — including by us?

  • On-device processing. Raw keystrokes, pointer paths, and voice never leave the device. Reduction to summary features happens locally, and the buffers are discarded.
  • Differential privacy on every output. Calibrated noise is added to every score before it leaves our infrastructure, so no individual can be reverse-engineered from any answer.
  • No raw data movement — architecturally. Our API has no endpoint that accepts raw behavioral events for storage. The guarantee is structural, not contractual.
  • Regulated from day one. NDPA 2023 and GAID 2025 at home, GDPR for global partners: explicit consent, data minimization, right to erasure — including a one-call baseline reset.

Common questions

Asked by every serious buyer.

Do you collect personally identifiable information?

No. The SDK never reads what a user types — only timing and motion patterns, which are reduced in the browser to a small set of timing statistics under opaque, device-salted labels. Device identity is a locally-generated random ID, hashed with SHA-256 before it leaves the device. There is no fingerprinting, no IMEI, no contact with the user's identity.

What happens if a legitimate customer gets flagged?

The verdict model is three-state by design: most uncertainty resolves to STEP_UP (ask for one more factor) rather than BLOCK. Low-risk sessions continuously refine the baseline, so a customer whose behavior drifts — new keyboard, injured hand — re-converges instead of being locked out.

How is this different from device fingerprinting?

Fingerprinting identifies a device across sites, usually without consent — it's a tracking technology. Lava Sense verifies that the person behind an already-authenticated session behaves like the account owner, inside one application, with consent, and with differential-privacy noise on every score. It answers "is this them?", never "who is this?".

All questions, answered

Design partners

The next piece of this
gets built with someone. You?

We're onboarding a small group of design partners — fraud, risk, data, and partnerships leaders at banks, telcos, insurers, and platforms.

Prefer email? Partnerships: partnerships@cleanlava.com