I translate between people who don’t share a vocabulary, then check whether the thing we built did what everyone assumed it would.
Both halves of that sentence came from the same four years, and I learned the second half because I got the first half wrong.
A trader told me the pricing screen was slow. The quant who owned the model agreed it was slow. So did the engineer who ran the service. Three people, one word, total agreement — and for most of a week we worked on three different problems. To the trader, slow meant the gap between wanting a price and having one, most of which was happening in his hands, not the system. To the quant it meant how often the model recalibrated. To the engineer it meant p99 latency on a service that was, by his numbers, comfortably fast. Nobody was wrong. Nobody was describing the same thing.
What fixed it wasn’t technical. It was getting the three of them to define the word in one room, and then measuring the thing they actually meant instead of the thing each had assumed the others meant. I’ve done some version of that every year since.
I spent four years as a product manager on FICC electronic trading at Bank of America, mostly on external vendor integrations and real-time data pipelines. Vendor work is the part I’d point at now: someone else’s system, someone else’s roadmap, your users’ deadline, and no authority over any of it. You get very good at finding the one question whose answer determines everything downstream, and at asking it before anyone has committed to a design.
It’s also where I stopped trusting stated requirements as a description of what people do. A desk that can’t opt out of your software will tell you immediately and unsentimentally when you’ve misread them. [Your turn: one or two sentences on why you left for the master’s — what you wanted to be able to build rather than specify.]
At Columbia I’ve been building the things I used to write specs for, and running into the same problem from the other side. On Sift, the event app I co-built and shipped, we put a taste questionnaire in front of every new user because a recommender with no signal can’t rank. Not one person finished it. Sixty percent of active users reached full personalization anyway, from swipes alone. We had assumed users needed to be asked. They’d been answering the whole time, and the only reason I could see it was that we’d instrumented both paths.
On Conviction, a retrieval system over SEC filings, I made the same mistake against a machine instead of a person. I built year-over-year risk-factor diffing on top of retrieval, because retrieval was the system I had. But top-k returns the passages most similar between two documents — which are precisely the ones that didn’t change. I was sampling boilerplate and diffing it. The feature only worked once I was willing to route around the thing I’d just finished building.
I’m also a research assistant at Columbia’s SEA Lab, working on a multi-agent system for mental rehearsal under Professor Xuhai Xu, with a paper under submission to CHI 2027.
[Your turn: one short paragraph that isn’t about work. Not a hobbies list — one specific thing you actually care about, in the voice you’d use out loud. This is the paragraph people remember.]
Conviction, and an MCP host called Scout that runs my own job search.
Multi-agent mental rehearsal at Columbia’s SEA Lab. Under submission to CHI 2027.
[Two or three things. Update when it stops being true, or delete this row.]
Full-time product and applied AI roles starting December 2026.