vinitium · case study

AI agents as first-class citizens

Role
Lead product designer
Duration
Multiple quarters, ongoing
Status
shipped
AI agents as first-class citizens — opening motif Many faint strands fan toward a gate; a chosen few continue on, drawn brighter — prediction as selection.

BrowserStack's test platform had capable AI agents that almost nobody used. I led the design initiative that reframed this as an experience problem — invisible AI is not adopted AI — and shipped the awareness system that changed it, then watched the numbers decay and got to work on why.

Problem

The platform's AI agents could heal broken tests, explain failures, and cut build times — and adoption ranged from near zero to roughly one in ten eligible teams. The capability was real; users either never discovered it, or enabled it and never saw evidence it was working. Two gaps, one root: the AI lived inside the walls, outside the user's line of sight.

Outcome

The awareness system shipped and the pattern was adopted by a sibling product team — validation that it generalized. Engagement climbed, then decayed over the following weeks. That decay, and the hypotheses we logged about it, are the honest ending: nudge fatigue is now the design problem, and it is not solved.

The setting

BrowserStack’s automation platform ships AI agents: one heals broken element locators mid-run (publicly reported to reduce automation build failures by 40%), one explains why a test failed in plain language, one trims test suites to only what a code change touches. These are genuinely useful capabilities with a real problem: inside the product, they were nearly invisible.

Adoption told the story. The failure-analysis agent — one click, in the debugging path everyone already walks — was used by nearly everyone. The agents that required configuration before their value appeared sat between near zero and roughly one-in-ten adoption. Same platform, same users, same quality of AI. The difference was never the model. It was the experience around it.

I framed the initiative around one sentence: invisible AI is not adopted AI. And a second, quieter one: value that isn’t demonstrated doesn’t retain. What follows is the decision that shaped the work — told as the three doors we stood in front of, including the two we didn’t walk through.

Door one: the AI hub — rejected

The obvious move. A top-level “AI” destination in the navigation: every agent’s insights in one place, weekly trends, recommended actions. A single pane of glass, and a strong demo.

We explored it seriously and killed it. The feedback that mattered, from research and internal review, was that users don’t visit destinations to appreciate AI — they live in their existing workflows and resent being asked to leave them. A dedicated page is real estate without footfall: a first-class nav item doesn’t guarantee first-class attention. The cons outweighed the pros, and the strongest argument for it (it would look impressive) was about us, not the user.

Killing this door taught the initiative its core principle: meet users where they already are, or don’t bother.

Door two: the floating button — deferred

A persistent, screen-independent affordance surfacing every agent with one-click enablement. High visibility by construction; a natural anchor for the chat-shaped agents on the horizon.

We deferred it rather than rejecting it. Too many unknowns at once: the agents’ individual workflows weren’t yet stable enough to hang off a single global affordance, and a persistent element competing with the product’s primary actions is a tax every screen pays. The note we left ourselves: revisit when the agent portfolio matures. Deferral with a written re-entry condition, not a polite no.

Door three: nudges in the workflow — shipped

The low-calorie option, deliberately: high-visibility cards and contextual prompts placed only on the highest-traffic surfaces — the pages users already visit dozens of times a day — each one dismissible, each one starting the agent’s enablement in a click or two rather than linking to documentation.

Anatomy of the awareness nudge pattern An abstract product dashboard. A nudge card sits inside the primary workflow surface, annotated with its three design choices: it lives on a high-traffic page rather than a separate destination, it is dismissible so it never hostages the workflow, and its call-to-action starts the capability in one click instead of linking to documentation. in the primary workflow, not a separate page dismissible — it never hostages the screen one-click start, not a docs link

The anatomy of the shipped pattern: in the primary workflow, dismissible, one-click start. Product surfaces abstracted by design.

The design fights were all about restraint. The dashboards these nudges live on are contested, crowded surfaces; product leadership’s constraint was non-negotiable — the user’s primary content stays above the fold. Every review pushed the same two directions at once: make it impossible to miss, and make it trivially easy to ignore. Holding both is the pattern.

What the numbers said

Indexed, by policy — the shapes are the story:

That last line mattered most to me as a designer: nobody who started got stuck. The funnel’s narrowness was an attention problem, not a usability problem — which is exactly what the door-three bet predicted.

The awareness funnel and its decay over time Left: a three-stage funnel where each stage is a small fraction of the one before — many saw the nudge, far fewer engaged with it, and a small share completed enablement. Right: a line of engagement over the weeks after launch — it climbs quickly, holds briefly, then declines steadily; the line ends as a dotted open question labeled nudge fatigue. saw the nudge engaged with it completed enablement launch weeks later nudge fatigue — the open problem

The funnel and its decay: a steep narrowing from exposure to enablement, then engagement climbing after launch, holding, and declining over the following weeks.

Then the honest part. Conversion degraded over time — early weeks performed multiples better than later ones. We logged hypotheses rather than excuses: nudge fatigue from repeat exposure, early saturation of the most-engaged cohort, seasonal effects, placement drift. The instrumentation to distinguish them shipped with the system.

What I learned about trust

Running structured reviews of the agents’ post-enablement experience (the method is its own artifact — see the companion essay on the scorecard) produced the sharpest finding of the initiative: our best agent scored well on clarity and relevance and failed on agency and trust. Users were shown AI-made changes as accomplished facts — nothing to accept, reject, edit, or gauge confidence by.

The four-pillar AI-UX review scorecard drawn as a web A radial chart shaped like a spider web with four spokes labeled clarity, agency, relevance, and trust. The plotted shape reaches far out on clarity and relevance but stays close to the center on agency and trust — the archetypal first review of an AI feature: it explains what it does, but gives the user little control and little reason to trust individual outputs. CLARITY AGENCY RELEVANCE TRUST the archetypal first review: legible, relevant — but the user can neither steer it nor gauge when to trust it

The archetypal first review: legible and relevant, weak on agency and trust — the profile that turns capable AI into unadopted AI.

Awareness gets a user to the door; agency and demonstrated value are what keep them in the room. The roadmap that came out of this — accept/reject/edit controls, confidence signals, explanations that say why this fix rather than what this feature does — is the initiative’s second act, in flight now.

Where it stands

Phase one is live in production. The pattern was picked up by a sibling product team for their own surface — the outcome I’m proudest of, because it means the thinking transferred, not just the pixels. Engagement decay is the open problem on my desk. If you’re fighting nudge fatigue in your own product, I’d genuinely like to compare notes.

The full walkthrough — screens, numbers, names — happens in conversation. Start one →