Independent data science practice

Machine learning that survives contact with production.

Krue helps early-stage product teams build personalisation, recommendation, and causal inference systems — from first model to something that runs, is measured, and holds up.

What I do

Four things, done properly.

Most early teams don't need an ML function. They need one specific system built well, and a clear read on whether it's actually working.

01
Personalisation & ranking

Recommendation, offer ranking, and content selection systems — including contextual bandits where exploration matters and cold start is real.

02
Causal inference

Uplift modelling and treatment effect estimation. Who to target, what the incremental effect actually is, and whether the lift is real.

03
Experiment design

A/B and multi-phase test design, power analysis, and the measurement layer that tells you whether a model earned its place.

04
First ML hire, fractionally

For teams pre-DS-function: setting up the data foundations, picking the right first problem, and building something that a future hire can inherit.

Selected work

Systems that shipped.

CONTEXTUAL BANDITS · MARKETING AUTOMATION PLATFORM
Offer ranking under uncertainty

Built a contextual bandit system for ranking offers across dozens of concurrent policies, using Thompson Sampling with a hybrid linear model. Included a live quality guard measuring ranking agreement in production, and a fallback chain so degraded policies never reached users.

+45% CTR
CAUSAL INFERENCE · QUICK COMMERCE
Who should actually get the notification

Designed an R-learner pipeline to estimate individual treatment effects for push campaigns across roughly 1.2M users, with a three-phase experimental design to validate targeting decisions. Found and diagnosed a feature leakage issue that had been inflating offline estimates.

~1.2M users · 10 campaigns
APPLIED LLMS · CONTENT GENERATION
Copy generation with a reward the business cares about

Reinforcement learning for notification copy, combining a quality judge, a conversion proxy, and a diversity penalty into a single reward. Ran and analysed competing training configurations to establish which reward weighting actually produced usable output.

SFT + GRPO

How it works

Scoped, paid, and small before it's large.

01
Conversation

Thirty minutes on what you're trying to decide or predict, what data exists, and what "working" would mean. Free, and often enough to tell you whether ML is even the right tool.

02
Scoping sprint

A short paid engagement under NDA. You get a written scope: feasibility, approach, effort, and the risks worth knowing before you commit. Credited against the build if we continue.

03
Build

The system, the evaluation around it, and the documentation your team needs to own it. Handover is part of the work, not an afterthought.

04
Measure

An experiment design that can actually detect the effect, so you know whether it worked rather than assuming it did.

Get in touch

Tell me what you're stuck on.

Early-stage teams, first ML systems, and problems where the measurement matters as much as the model. Happy to sign an NDA before you share anything.

mansi@krueai.in