At a glance
- AI-written explanations shipped, up from a few thousand human-checked lines
- 157,000 a day
- agreement between the AI graders and human reviewers, and the graders came out stricter
- 96 to 100%
- incidents we had to walk back, shipping AI at catalog scale
- 0
Netflix had spent two decades learning what people watched. What they finished. What they rated. The shows they hovered on for a second longer than everything around them. It knew an enormous amount, and it rarely explained any of it.
- My role
- Senior Product Designer, personalization and explanations, Netflix member experience. I led the design. A PM owned prioritization, engineering and data science owned the model plumbing.
- Timeframe
- 2021 to 2026 · three moves, then a three-year north star
- Where it shipped
- TV, mobile, and web. Detail pages, rows, billboards, post play.
- The one thing
- I replaced a match score nobody believed with AI-written reasons a title is for you, and designed the grading system that let them run at about 157,000 a day with zero incidents walked back.
The system knew why. It never said it.
Fewer than half of members felt their home page was personalized to them. For the company that invented personalization. I could not get past that number. We had a system making good decisions, and from the member's side it often looked like a wall of recommendations.
The closest thing we had was a row called Highly Recommended For You. It tells you how confident we are, and nothing about you. So the member's real question stayed unanswered. Why this one, and why should I care.
The engine knew why. The member saw a poster and a promise.
The bar was a friend, not a label
I dug into the research, and our instinct had been backwards. When members struggled to choose, we added something. Another badge. Another piece of information. Another thing on the box. People kept asking for something simpler. Why this one?
One study stayed with me. A woman told us she had avoided Stranger Things for years because we labeled it horror. She only tried it after a friend she trusted recommended it. That was the standard I carried into everything after: not a label, a friend.
HorrorFour years, in stages. Each step got closer to the person.
I spent the next four years on this in stages. Each version got a little closer to the person. The sentence took an afternoon. Getting to show it took years.
First, give people control. Then, an honest number. Then, a reason built on what they did. Then, a reason written for them. Each step asked the same thing: what have we earned the right to say?
Try one: the wheel in their hands
My first instinct was straightforward. If we were not going to explain the recommendation, maybe we could at least give people more control over it. So I built controls, taps and filters that let you steer your own recommendations, at different points in the journey: browsing the home page, after you finished something, on the title itself.
Some of it worked. Rows that re-sorted around what you were in the mood for got a real lift. In one study a woman typed THANK YOU NETFLIX in all caps when a row finally moved the way she wanted.
But control turned out to be a completely different thing from understanding. I had given her a steering wheel without saying where the car was going. She still had no idea why.
The number I stopped trusting
Before I could add a better reason, I had to deal with a bad one. For years Netflix showed a match score. 92% match for you. It sounded precise. I kept noticing it and not believing it, and I started wondering whether a number like that even made sense as evidence.
So I ran a holdback. Turned the score off for one large group of members, left it on for another, and let the metrics talk. Nothing happened. Not one metric moved. It had sat on our titles for years looking scientific and doing nothing.
A member in research put words to why it was worse than nothing. If Netflix says highly recommended and I hate it, fine, they got one wrong. If Netflix says 92% and I hate it, now I wonder what is wrong with me. Our fake precision took our own miss and made her feel broken.
The lesson I carry into every AI product since is simple. A confident number you have not earned costs more trust than saying nothing at all.

A confident number you haven't earned costs you more trust than saying nothing at all.
Try two: an honest line, then a real reason
So I killed the score. In its place we tried an honest line built on the member's actual thumbs ratings: Highly Recommended For You, based on what you've watched and rated. It sat next to shows they had genuinely rated. I built it with my content designer and our engineers, and the honest version beat the impressive-looking one. More people rated. Streaming went up. Small, but it gave us a direction.
Still, that line was our confidence with a nicer face on it. The next step was to give a reason drawn from what you had watched. Because you watch crime thrillers. Because you watch shows set in the eighties. Because you watch movies based on books. You could look at one and think, okay, I know why you are saying that. Members felt a little more seen.


A writing problem and a care problem
Now it was a writing problem and a care problem. The work was the rules: how to describe a title with taste, so a line lands like a good recommendation and not a sales pitch. We could describe what someone watches. We could not make claims about who they are. We never called anyone a fan, because the moment you label a person you get personal fast. We also screened out anything that could sting. Nobody should get a little Netflix message that says, because you watch shows about addiction.
Describe what someone watches. Never claim who they are.
A category still is not you
People responded. They felt more seen than a number ever made them feel. Then I started using it myself and hit the same wall. A category still is not me. Sure, I watch crime thrillers. But why this one? Why tonight? What is it about this exact title that connects to the thing I loved in something else?

Meanwhile, at home
I already knew what a deeper version could feel like, because I was doing it at home. I picked shows two ways: friends told me what they loved, and I asked a large language model. I would describe exactly what hooked me, a theme, an actor, the mood I was in, one specific turn in the story, and ask what to watch next. It got surprisingly good. Eventually I realized I had built myself a recommender.
So I brought that instinct back to Netflix and pushed for it: a reason written by a language model that names the real, personal thread between your taste and one specific title.
▋
Try three: two things read at once
In 2024 I led the design of the first time Netflix put AI-written words directly in front of members.
Roughly, the model reads two things. Everything we know about you: what you finished, what you rated, the themes you keep returning to. And everything about the title: its cast, its tone, the shape of its story. Then it finds the actual thread between the two and says it in one human sentence. The kind of reason a friend would give you, instead of a score.
the themes you keep returning to
the shape of the story

Humans could not check the lines fast enough
The writing was new. Sometimes the model made things up. Sometimes it paired two titles that had no business together. Every line had to be checked by a human, and a human cannot keep up with a catalog that size. We were drowning in review.
A grader, not a writer.
A second AI grades the first
The fix is a design idea more than a technical one. Quality checking is usually the boring step at the end. I moved it to the middle and built the product around it.
With our data scientists I set up a second AI whose only job is to grade the first. I designed the three checks and the rule for what happens on a fail. They built the models, and we tuned the rules together. Three checks. Is it true? Is it sensitive in a way we should avoid? Do these two titles belong together? That third one worried me most. A strange pairing is harmless in front of one researcher. In front of a few hundred million people the stakes change. So that judge got its own dedicated model.
Clear the bar or vanish
The rule I fought hardest for is simple. If a line cannot clear the bar, we do not show a weaker version. We show nothing. That sounds conservative, and it is, and it is why it works. The only reason a member trusts the good explanations is that the system is willing to say nothing.
96 to 100% agreement with humans. 157,000 lines a day. None walked back.
We checked the AI graders against our human reviewers. They agreed 96 to 100 percent of the time, and came out stricter than the humans, which is the right direction for a grader to be wrong in.
That one move took us from a few thousand human-checked lines to about 157,000 explanations a day. Real watching went up. Zero incidents we had to walk back. Bigger and safer at the same time.
Then it taught me one more thing. On mobile, on the detail page, the explanation helped while someone was still deciding. On TV the home screen had already done most of the persuading, and that same paragraph one click from play made people stop and read instead of pressing play. Nothing about the words changed. What changed was where in the decision they showed up.


When the bottleneck is a human reading every line, the evaluation system is the product.
The North Star: one week in Sam's life
I felt the company was not moving fast enough on AI, so I did not wait for permission. I pulled together a small tiger team, a lead engineer, a data scientist, a content designer, and a PM, and led the design of a three-year north star.
Then I walked leadership through it the way I designed it. Not a parade of features. One made-up member, Sam, 34, East Bay, who binged Squid Game, finished Ozark, and stalled on Wednesday at episode six. One week of his life, every idea a moment in his week. The core was to stop hinting and name the reason, and let the explanation travel with Sam everywhere: the home page, a hover, the detail page, even the credits. Recommendation and explanation as two outputs of one engine.
Monday, he picks up where he left off, and one sentence tells him what he has to gain by going back. Tuesday, the same title is explained three ways, each tied to something he already loves, never crossing into misrepresentation. We did not tell him what it is. We told him what it is to him. Wednesday, a cold start: he just says "a dark, complex show" and steers by mood. Thursday, a podcast framed as the natural next step from a show he loved. Friday, a game, a podcast, and a film tied by one thread he has expressed for years. During playback, a light guide says who this character is and what this song is, without pulling him out. Saturday, three explained next watches over the credits. Sunday at 10:30, alone with work tomorrow, one standalone episode, done by midnight.
I put a whole slide on the brain behind it, the judges, the audit, the fallbacks, because beautiful writing with no quality system underneath is AI slop at scale. And I refused vanity metrics and led with coverage, what percent of titles even get an explanation. Two good reasons out of eight titles feels personal. Eight out of eight feels like noise. If Netflix tells you you'll love everything, you stop believing any of it.
A boy’s disappearance exposes sinister secrets and terrifying creatures in this small-town mystery that echoes the creeping dread of “Hill House.”
Packed with thrills and ’80s nostalgia, this story of friends uncovering a hidden world became a global hit, sparking a wave of viral energy like “Wednesday.”
A group of kids team up to face monsters from another realm in this thrilling adventure full of bike chases, spooky mysteries, and strong friendships.







One pipeline. One generation logic. Every surface.
How it landed, honestly
I will be straight about how it landed. It won the room and changed the roadmap conversation. Full funding was still competing with a big platform rebuild when I left.
What I took away is how hard it is to push an AI-first idea inside a company not yet built around AI. You can have the research, the prototype, the numbers, and the vision, and still run into the shape of the organization. That is a big part of why I want to be somewhere that is built around it.
- Explicit controls
- Audience explanations
- GenAI explanations
Four years, one step at a time
First we gave people some control. Then we got rid of a number that looked precise and did not help. Then we grounded the explanation in things a person had genuinely done. Then we let AI find a deeper connection, and built a system around it that could decide when the explanation was not good enough.
All of those were trust decisions. Can I stand behind this number? Can I prove this reason came from something the member actually did? Can I put this explanation in front of someone without an assumption I have not earned? And most of all, can the system admit when it does not know?
It is hard to get someone to trust a small signal enough to act on it. AI makes it bigger, because the system can say so much, so quickly, with so much confidence. I spent years figuring out when Netflix should speak, what it had earned the right to say, and when it was better to leave the space empty. An assistant faces that call on every turn. Answer, hedge, or say nothing. The grader was my first version of it at scale. It is the work I want to keep doing.