At a glance
- badge coverage on TV Home, grown on purpose, then capped
- 8.3% to 24.5%
- incremental value from one badge test
- $8M a year
- share of people who pressed play, from the Top 10 badge alone
- +11 points
Here is the exact moment I designed for. It is a Tuesday, 9:40 at night. You have made a thousand tiny decisions today and have none left. You want something good, and you have about two minutes before you give up, put The Office back on, or close Netflix for YouTube.
Hundreds of thousands of titles. About forty rows on your screen. That is the whole job: help you decide fast, and not regret it thirty seconds in. Getting you to a choice fast is half of it. The other half is trust.
- My role
- Lead designer on the evidence framework for the Netflix member experience. I owned the text evidence layer end to end, across TV, mobile, and web, including the evidence layer of the video forward TV redesign.
- Timeframe
- 2021 to 2026 · five years
- Where it shipped
- TV, mobile, and web. Games, live, ads, series, and films all came through the one framework I owned.
- The one thing
- I turned four overlapping evidence systems into one governed framework: what a title is allowed to say, how much, and when, ranked in code. Coverage grew from 8.3% to 24.5% and I capped it there.
Evidence: every true thing a title says to help you decide
The thing I owned is what we call evidence. All the little pieces of information around a title that help you decide. Number one this week. The final season. The actor you love. Based on the book you read.
Every piece has to be true. We are handing you real reasons, so when you press play, you trust we were being straight with you.
That line under the art, "Spent 6 weeks in the Top 10," is a piece of evidence. So is the trophy on the poster, or the "New Episode" badge. Every one true, every one earning its place on the art.

Every team came through one framework
I led the design of that system for the entire member experience org. I was the person every other team came to when they wanted to say something on a title. Games, live, ads, series, films. They all came through me and the framework I owned.
The mess: four systems, ten years, nobody coordinating
When I got there it was a mess, and a very human one. For about ten years, team after team had bolted their own badge or line of text onto titles. The same title might say four different things in four different places, and the loudest message usually won, not the most important one.
The home page had turned bright red and shouty. We had a nickname for it, the NASCAR effect. A race car with so many sponsor stickers you cannot tell what car it is. When everything is shouting, you cannot hear any of it. That is a person we lost, at 9:40 at night, because we could not get out of our own way.
Think of the product as a sand castle, and everybody pours their own pile of sand in. Something inevitably slips. My job is to keep the castle standing while it keeps growing.

Every message the system could show
So I did the boring thing. I catalogued every message the system could show, across all four old systems, into one enormous spreadsheet. Hundreds of them. The same idea said five ways by five teams who never talked.










You cannot fix a mess you cannot see.
The argument I had to win, and the receipt I brought
This was the hardest part of the job. Every single team wanted more. More badges, more messages, more of their thing on the title. And every one had a good reason. From where they sat, they were right.
I could not win that by saying no louder than them, and I could not win it by letting everyone pile their stuff next to each other. So instead of arguing, I forced conversations, and I brought receipts to every one of them.
From where they sat, they were right.
The eye-tracking studies
My first receipt was the research. I pulled the eye-tracking studies. People are barely reading. They skim on instinct, what researchers call system one. Every so often something catches them and they slow down and read, that is system two. So I would sit a partner team down and say, here is the moment you are fighting for. You do not get a paragraph. You get a glance.





Progressive disclosure: the evidence earns more words as you lean in
At Uber I had designed scooter unlock to reveal each step just in time. Same instinct, grown up. At Uber it was about sequence. Here it is about attention, matching how much the evidence says to how much you are giving it in that moment.
On the home page, skimming, you get the single strongest, most personalized reason and nothing else. One thing, the best thing. Hover or focus on a title and you have told us you are a little curious, so the evidence earns the right to say a bit more. Click into the detail page, where you are weighing whether to press play, and that is where the fuller story lives, because now you are reading.
Deciding when to speak, and how much to say, based on what a person wants in the moment, is the whole job of an assistant too. Same design problem, one surface over.






Taste does not scale, so restraint became a structure.
The hierarchy and the ceiling
To make it stick across an org that size, I built a hierarchy. A stack rank of all the evidence types. Eleven categories, each ranked against the others, enforced in the code, so the hierarchy was the authority, not whoever argued hardest that week.
I leaned on research to split table stakes, what a member has to see, like whether they can even watch it right now, from nice to have, like a trivia fact. Table stakes gets the slot. Nice to have waits for you to lean in.
And we grew badge coverage on TV Home from 8.3% to 24.5% of impressions, then capped it on purpose.
- 01Unique Content Cases
- 02Timeliness and Freshness
- 03Member Intent
- 04Social Proof
- 05Social Proof plus Thumbs
- 07Viewing History
- 08Need State
- 09Awards and Praise
- 10Previously Live
- 11Talent
- 12Title Features

The hierarchy was the authority, not whoever argued hardest that week.
The test that proved it
A principle on a slide is cheap. I ran a test. The old way let the strongest few categories win outright, so a title often said the same kind of thing twice, two flavors of popular, and the second line wasted the slot. I tested a rule: the stack could take only one callout per theme before it had to move down to the next one. Suddenly the second and third things a title said each added something new.



Sixteen cards became four. Four systems became one. On the most watched screen in the company.
Around 2024 Netflix did its first real redesign of the TV experience in about a decade. It made the news. A research-backed bet on full video instead of a wall of tiles. A single point of focus. I was one of the core designers on it. I owned the text evidence layer, not the whole redesign.
My stomach was in knots. Going video forward meant even less room for evidence than before. Less space, higher stakes, on TV Home. Every decision above, the consolidation, the ranking, the progressive disclosure, got harder and mattered more. And I was carrying the whole evidence system onto a new platform while that platform was still being built underneath me.
So I bet on stripping evidence down to its most glanceable, most essential form, and finally folding that ten-year-old pile of disconnected systems into one clean framework. Four legacy types, badges, supplemental messages, evidence cards, and tags, became one callout framework. Sixteen evidence cards at sixty characters came down to four at forty or fewer.
10#1 in TV Shows👍 Most Liked

Like re-plumbing a house while people are moving in.
Quiet, stubborn iteration
Getting there was not one clean move. I would design a stripped-down version. We would drop it into that shrinking, video forward frame, look at it, and realize it was still too much. So we would cut again. Me, my content designer, and the engineers, arguing over single words and single pixels, until what was left was only the things that earned their place on that screen. That narrowing down is the actual craft.
One system, tuned to every surface
Evidence did not live on one screen. The same system had to hold across TV, mobile, and web, and it was shared by games, live, ads, series, and films.
Each surface is a different moment. On TV the home is evidence rich, so you can explain a little later, once someone leans in. On mobile the home is almost all cover art, so the evidence has to show up earlier, and smaller. The rules held on every surface. The layout did not have to. I wanted the underlying system to hold everywhere without forcing every surface into the same shape.
Availability: a question Netflix never had to answer until live
This is the one I love most, and it is the least glamorous, which is very on brand for me. Availability evidence answers one simple question: when can I watch this?
For most of Netflix's history that question barely existed. Everything dropped at once. Then came things with a clock on them, live sports, weekly episodes, games, and suddenly "when" was a real product question. It reordered the questions I designed around. What is this. When can I watch this. How do I watch it. And only then, why should I care.
At the time all that timing information was jammed into the same badge system as everything else, right next to "number one this week" and "new season." But availability does a different job. A promotional badge can rotate in and out. Availability has to be there the moment it matters. Treated like a promo, it kept getting bumped by louder messages.
So I made the case to pull it into its own evidence type, with its own component, rules, and priority. Anything that affects whether you can watch, not out yet, expired, not on your plan, now had a dedicated place and could not get buried. It sits below the title on purpose: research found higher comprehension when availability is paired with title-level information.







The lifecycle: Coming Sunday, Live now, Started at, Leaving soon, Gone
Then came live, the hardest version of it. When we launched NFL games I did not control the timing. It came from the rights deal, and it was unforgiving. One game might be watchable for hours after it aired. The next for minutes. Outside the US, a different window again. So the same title had to say something completely different depending on the exact minute you looked.
I designed that entire lifecycle, every state a live event moves through and what you should see in each one, and it had to be right on every screen at the same time. We tested it on real NFL Sundays, watching an actual game move through those states in production. "Leaving soon" had to appear while you could still catch it. "Expired" could not show a second too early. Kickoff is kickoff.
The first launches were hand built, every non-standard line an engineer typing a one-off override. That would not scale as Netflix did more live, so I designed the availability component to handle the logic itself. It figures out whether something is daily, weekly, or a one-time drop, and it knows what to say when two things are true at once, like a show that is both live and weekly. Any team can use it without a designer working out the same logic from scratch.


So it could run without me: docs, courses, and a chatbot on the docs
By this point so much of the org was using the framework that it had to work without me. So I wrote the whole system down: the hierarchy, the categories, the priority rules, composition, surface differences, edge cases, and the reasoning behind them. I ran office hours. I built short courses so designers, PMs, engineers, and partner teams could learn it. And I put a chatbot on top of the documentation to answer routine questions. Then it was called NotebookLM; now it is called Gemini Notebook.
The point was to let the framework scale without turning me into its permanent help desk.












$8M from one test, +11 points from one badge, five years coherent
It paid off in the numbers Netflix cares about. One badge test I ran was worth about $8M a year. The strongest single badge, Top 10, lifted the share of people who pressed play by about eleven points. And the system stayed coherent through five years and a half dozen new content types.
The numbers are what I can show you. What I took from it is this. Helping a member trust a small signal, so they decide whether to watch, is the same shape of problem as helping a person trust an AI's answer, so they decide whether to act on it. Same job, one layer up.
That is what pulled me toward AI
Because the system had gotten that good, the one thing it still could not do started to bother me. It was still talking to everybody the same way. Netflix knew an extraordinary amount about what each member liked, and almost none of it ever made it onto the screen.
So the next question became: can the evidence tell you why this title might be right for you? That is what pulled me toward AI.




