vinitium · case study
AI agents as first-class citizens
BrowserStack's test platform had capable AI agents that almost nobody used. I led the design initiative that reframed this as an experience problem — invisible AI is not adopted AI — and shipped the awareness system that changed it, then watched the numbers decay and got to work on why.
Problem
The platform's AI agents could heal broken tests, explain failures, and cut build times — and adoption ranged from near zero to roughly one in ten eligible teams. The capability was real; users either never discovered it, or enabled it and never saw evidence it was working. Two gaps, one root: the AI lived inside the walls, outside the user's line of sight.
Outcome
The awareness system shipped and the pattern was adopted by a sibling product team — validation that it generalized. Engagement climbed, then decayed over the following weeks. That decay, and the hypotheses we logged about it, are the honest ending: nudge fatigue is now the design problem, and it is not solved.
The setting
BrowserStack’s automation platform ships AI agents: one heals broken element locators mid-run (publicly reported to reduce automation build failures by 40%), one explains why a test failed in plain language, one trims test suites to only what a code change touches. These are genuinely useful capabilities with a real problem: inside the product, they were nearly invisible.
Adoption told the story. The failure-analysis agent — one click, in the debugging path everyone already walks — was used by nearly everyone. The agents that required configuration before their value appeared sat between near zero and roughly one-in-ten adoption. Same platform, same users, same quality of AI. The difference was never the model. It was the experience around it.
I framed the initiative around one sentence: invisible AI is not adopted AI. And a second, quieter one: value that isn’t demonstrated doesn’t retain. What follows is the decision that shaped the work — told as the three doors we stood in front of, including the two we didn’t walk through.
Door one: the AI hub — rejected
The obvious move. A top-level “AI” destination in the navigation: every agent’s insights in one place, weekly trends, recommended actions. A single pane of glass, and a strong demo.
We explored it seriously and killed it. The feedback that mattered, from research and internal review, was that users don’t visit destinations to appreciate AI — they live in their existing workflows and resent being asked to leave them. A dedicated page is real estate without footfall: a first-class nav item doesn’t guarantee first-class attention. The cons outweighed the pros, and the strongest argument for it (it would look impressive) was about us, not the user.
Killing this door taught the initiative its core principle: meet users where they already are, or don’t bother.
Door two: the floating button — deferred
A persistent, screen-independent affordance surfacing every agent with one-click enablement. High visibility by construction; a natural anchor for the chat-shaped agents on the horizon.
We deferred it rather than rejecting it. Too many unknowns at once: the agents’ individual workflows weren’t yet stable enough to hang off a single global affordance, and a persistent element competing with the product’s primary actions is a tax every screen pays. The note we left ourselves: revisit when the agent portfolio matures. Deferral with a written re-entry condition, not a polite no.
Door three: nudges in the workflow — shipped
The low-calorie option, deliberately: high-visibility cards and contextual prompts placed only on the highest-traffic surfaces — the pages users already visit dozens of times a day — each one dismissible, each one starting the agent’s enablement in a click or two rather than linking to documentation.
The anatomy of the shipped pattern: in the primary workflow, dismissible, one-click start. Product surfaces abstracted by design.
The design fights were all about restraint. The dashboards these nudges live on are contested, crowded surfaces; product leadership’s constraint was non-negotiable — the user’s primary content stays above the fold. Every review pushed the same two directions at once: make it impossible to miss, and make it trivially easy to ignore. Holding both is the pattern.
What the numbers said
Indexed, by policy — the shapes are the story:
- Tens of thousands of users saw the nudges in the first months.
- Roughly 7 in 100 who saw one engaged with it.
- Of those who engaged, about 6 in 10 carried through toward enablement.
- The median time between steps was seconds — once users committed, the flow itself was frictionless.
That last line mattered most to me as a designer: nobody who started got stuck. The funnel’s narrowness was an attention problem, not a usability problem — which is exactly what the door-three bet predicted.
The funnel and its decay: a steep narrowing from exposure to enablement, then engagement climbing after launch, holding, and declining over the following weeks.
Then the honest part. Conversion degraded over time — early weeks performed multiples better than later ones. We logged hypotheses rather than excuses: nudge fatigue from repeat exposure, early saturation of the most-engaged cohort, seasonal effects, placement drift. The instrumentation to distinguish them shipped with the system.
What I learned about trust
Running structured reviews of the agents’ post-enablement experience (the method is its own artifact — see the companion essay on the scorecard) produced the sharpest finding of the initiative: our best agent scored well on clarity and relevance and failed on agency and trust. Users were shown AI-made changes as accomplished facts — nothing to accept, reject, edit, or gauge confidence by.
The archetypal first review: legible and relevant, weak on agency and trust — the profile that turns capable AI into unadopted AI.
Awareness gets a user to the door; agency and demonstrated value are what keep them in the room. The roadmap that came out of this — accept/reject/edit controls, confidence signals, explanations that say why this fix rather than what this feature does — is the initiative’s second act, in flight now.
Where it stands
Phase one is live in production. The pattern was picked up by a sibling product team for their own surface — the outcome I’m proudest of, because it means the thinking transferred, not just the pixels. Engagement decay is the open problem on my desk. If you’re fighting nudge fatigue in your own product, I’d genuinely like to compare notes.
The full walkthrough — screens, numbers, names — happens in conversation. Start one →