Launch-ready is a claim you have to earn.

I work with founders and with the engineers who own the code, on what the product does today, what it can claim in market, and what has to happen before either is true.

product intermediator applied ML engineer responsible AI AI evaluation voice AI auditor technical author @W&B community builder
Shreyan Basu Ray Shreyan Basu Ray, photographed

that's me, Shreyan ✦

17+
technical articles for weights & biases ✦ read them →
820+
builders in claude developers india ✦ ask to join →
1
published book ✦ The Architect's Handbook →

What I've shipped

all cases →
enterprise voice agent platform ✦ in production

A voice agent that runs at ₹2.29 a minute

Built as backend infrastructure so the client can stand up voice AI customer support for any company that needs it. I costed it layer by layer before writing code, and the stack was chosen against that model rather than against a benchmark.

₹2.29 per minute, all in, at production quality. It is live and running.

per-minute variable cost
STT, Google standard$0.024
LLM, Gemini 2.0 Flash$0.0001125
TTS, Neural2 / HD$0.000008
RAG, retrieval + embeddings$0.00002
total $0.02414

₹2.29 at ₹95/USD

risk intelligence ✦ fintech

The metric was measuring the wrong thing

The reported distress rate was climbing, 21% to 27% year over year, and it read like a market signal. It was 187 spurious labels from companies that had simply not filed their accounts. I traced it to the label definition rather than to the model, and added a reportability gate that holds the rate inside a point.

92.155% on current-year distress flagging ✦ SHAP attached to every prediction

read in detail →
natural language to SQL ✦ facilities SaaS

₹95,000 a month for bad answers

A facilities management platform was paying a vendor copilot to answer questions about their own data badly. I modelled the unit economics layer by layer before writing code: the cheapest model for intent routing, semantic table selection so only relevant schemas reach the LLM, three tiers of cache.

₹95,000 to ₹35,000 a month ✦ post-deployment spend tracked the model

read in detail →
independent evaluation ✦ AI coding agent

An outside pass before it shipped

An AI coding agent in private alpha brought me in to evaluate the product before wider release. I ran it across six tracks, from static review and security pentest to LLM red-teaming and capability evaluation, each scored against rubrics the industry already recognises, from OWASP LLM Top 10 to CWE, MITRE ATLAS and SWE-bench.

18 findings, ranked and reproducible ✦ the attacks it was built to stop held, the gaps were architectural

read in detail →

Partners I bring in

✦ marketing · tech · legal

Specialists I trust for the parts beyond me. I own it for you, so your product ships ready for the market at the standard it deserves.

✦ my partners work end to end under my advisory brief, and I take care of the outcome

how partnering works →
where I fit

Who I work with

Where a company sits on this scale decides what I can do for it. I go deepest with small and mid-sized businesses, close enough to evaluate the product, diagnose what is holding it back, and settle what it can honestly claim in market.

only when structured
Very small enterprises
1 to 9 people

I take these on when the team is highly competent or already structured, where being small is an edge rather than a risk.

main ground ✦ deepest here
Small and mid-sized businesses
10 to 100 people

My main ground, and where post-seed startups at Series A and B sit too. I evaluate the product, diagnose what is holding it back, and settle what it can honestly take to market.

evaluation and strategy
Mid-market
100 to 1,000 people

I help them evaluate the product, diagnose how it is working and where it is heading, and strategise the GTM around their positioning and where they can win customers.

a different game
Enterprise
1,000+ people

Enterprise runs on field sales, procurement and multi-year cycles. That is a different discipline from the fast-moving work I am built for.

← smallercompany sizelarger →
How I consult →

Built for the Opensource

github →

cc-habits ↗

open source

developers across nine coding tools

A tool-agnostic memory layer that learns a developer's habits from real edits and writes them into whichever coding agent they are using, so your context follows you between tools instead of starting cold every session. Hardened against symlink races, Unicode sanitiser bypass and indirect prompt injection.

install from npm ✦ security audit public in the repo

A hands-on path through quantum machine learning that is honest about where it helps and where it does not, including problems set so the expected answer is that the quantum method loses to a classical one. Built to learn from, not to sell quantum.

MIT ✦ running since december 2024

A study of how well text watermarks survive once content is paraphrased and regenerated, so a team can judge whether a watermark is worth relying on. Carries an explicit responsible-use note citing EU AI Act Article 50(2).

Research I've published

google scholar →
working paper ✦ july 2026

Agent Complexity Theory

Choose an agent architecture before you build it, not after benchmarking.

A computational theory that characterises an agentic workflow by the intrinsic difficulty of the task and the complexity signature of the architecture. Signatures derived for four machine families with matching lower bounds, and reconciled against ten production workflows.

33 pages ✦ four complete case studies ✦ one documented failure case

read the paper →

framework

SafeTrust: Agentic Development in India ↗

A framework for building and deploying agentic systems responsibly in the Indian context.

peer reviewed

Quantum Neural Networks ↗

A book chapter on where quantum machine learning actually earns its cost against classical baselines.

Springer Nature, with Soujanya Ray ✦ ISEM vol 65, pages 205 to 233 ✦ first online september 2025

book

The Architect's Handbook: Engineering Responsible AI ↗

A practitioner's handbook on building AI systems that can be governed, audited and explained.

Eighty-six pages, free to read, on Kindle ✦ version 1.0, january 2026

policy

C-RAI ↗

A certification framework for responsible AI, proposing how Indian AI systems could be audited and labelled before deployment.

Sent to MeitY and NASSCOM ✦ policy draft v2, 2026

Articles I've written

all 17 →
Getting started with embodied chain-of-thought (ECoT) weights & biases ✦ jun 2026 GLM-5.2: The open coding model that put the critic back into RL weights & biases Fine-tuning Qwen 3.5 on multimodal data with Unsloth and W&B Weave weights & biases Understanding the Muon optimizer, tested against AdamW with NanoGPT weights & biases

I run every experiment myself before writing it up.

Agent Complexity Theory ✦ cc-habits ✦ Physical AI ✦ Claude Developers India ✦ C-RAI ✦

The dev community

claude developers india ✦ 820+ builders

I founded Claude Developers India, a WhatsApp community of 820+ builders, and I keep it active. I run "Show Your Products, Brainstorm and Validate", online calls where members demo what they are building to each other and get feedback. Each session also brings in a guest speaker or two for a short talk.

I also mentor at hackathons, most recently Code Wizards 2.0 at SRM University, Ghaziabad Campus.

Ask to join on WhatsApp → Upcoming sessions
Claude Developers India meetup group photo

the crew, after Claude event ✎

In conversation at a NASSCOM event

at NASSCOM ✎

Helping a participant at a laptop

one bug, two heads ✎

Mentor board at Code Wizards 2.0

mentoring at Code Wizards 2.0 ✎

Community meetup selfie

name tags on, still talking ✎

Two people deep in conversation across a cafe table

the hallway track ✎

Want an outside pass over your AI product?

Three questions and I will tell you whether it is worth a call.

Answer them below ↓
or email hello@shreyanbasuray.com
agency or studio ? Be a Partner →
what a silent failure looks like voice agent ✦ one turn
okstt.transcribeaudio in
okintent.routematched
failrag.retrieve0 rows returned
okllm.generateanswered anyway
oktts.speakspoken to caller

no crash. no error raised. nothing in the dashboard.
the caller just gets a confident wrong answer.

start here ✦ two minutes

Tell me what you need

Pick the closest fit. The questions change with it, and none of them take long. It goes straight to my inbox, nothing else.

an AI product going to alpha or launch, and you want an outside pass before it does

github linkedin x scholar weights & biases npm events whatsapp email
reviewed october 2026
WhatsApp me