Designing an AI Agent developers could actually trust with their APIs

Designing an AI Agent developers could actually trust with their APIs

Designing an AI Agent developers could actually trust with their APIs

I designed API Agent to turn complex, multi-step API workflows into a guided 15-minute experience using AI to do the heavy lifting while keeping developers in control.




I designed API Agent to turn complex, multi-step API workflows into a guided 15-minute experience using AI to do the heavy lifting while keeping developers in control.




Role:

Primary designer

Team

UX researcher, PM, Dev

Timeline:

2025 - 2026

~15 min

~15 min

~15 min

to create an API with Agent, down from ~3 days manually

47.8K

47.8K

47.8K

monthly queries, +23% month-over-month

78%

78%

78%

of sessions ended in a successful outcome

CONTEXT

The problem wasn't capability. It was complexity.

Developers could create, test, troubleshoot, and manage APIs in API Connect BUT these workflows involved many steps, configurations, and tools strung together. Early research showed that manually creating an API could take roughly two weeks. AI presented an opportunity to rethink how developers moved through this complexity—not simply add another chatbot.

How might we design an AI Agent that turns complex API workflows into simpler, intent-driven interactions while keeping developers informed and in control?

RESEARCH INSIGHT

Developers wanted assistance, not autonomy.

Developers wanted assistance, not autonomy.

Through customer interviews, advisory board sessions, usability testing, and private preview feedback, we explored where AI could create value across the API lifecycle. Developers wanted speed, automation, and better guidance but the strongest signal wasn’t about capability. It was about trust.

Through customer interviews, advisory board sessions, usability testing, and private preview feedback, we explored where AI could create value across the API lifecycle. Developers wanted speed, automation, and better guidance but the strongest signal wasn’t about capability. It was about trust.

Discover

Create

Secure

Test

Manage

Socialize

“If I trusted it, it’d be fantastic but I would distrust this from the get-go.”

“If I trusted it, it’d be fantastic but I would distrust this from the get-go.”

Developer, usability session

“I don’t want AI to write code for me that I don’t understand.”

“I don’t want AI to write code for me that I don’t understand.”

Developer, usability session

DEVELOPERS WANTED

BUT WORRIED ABOUT

Faster delivery

Faster delivery

Loss of control

Loss of control

Smarter automation

Smarter automation

Hallucinations

Hallucinations

Generated code

Generated code

Not understanding what was generated

Not understanding what was generated

AI recommendations

AI recommendations

Whether recommendations were correct

Whether recommendations were correct

Less manual work

Less manual work

Approving changes blindly

Approving changes blindly

This tension showed up consistently across creation, discovery, testing, governance, troubleshooting, and security work. Rather than designing around productivity alone, I designed around trust directly.

This tension showed up consistently across creation, discovery, testing, governance, troubleshooting, and security work. Rather than designing around productivity alone, I designed around trust directly.

HOW I WORKED WITH RESEARCH

DESIGN DECISIONS

01.

Plans make actions predictable

Plans make actions predictable

Too much confirmation created friction. Too little undermined trust. So I split Agent actions into two types: read actions, like search and inspect, run immediately since they carry no risk. Write actions, like create and deploy, show a plan first, so developers always know what's about to happen before it does.





02.

Traces make progress visible

Traces make progress visible

Once the Agent started working, silence created uncertainty. I designed a lightweight progress trace to show what’s happening, what’s done, and whether it’s still working without exposing every backend step.






03.

The UI makes results actionable

The UI makes results actionable

I separated explanation from action: chat explains, the UI applies.


When a developer asks the Agent to fix an error, the change is applied directly in the editor and highlighted with a brief gradient that fades after it’s been seen. This gives developers clear proof of what changed without adding permanent clutter.


The pattern became the foundation for subsequent Agent experiences.





08 — FROM FEATURE TO SYSTEM

The flagship wasn’t the feature. The interaction model was.

API creation was the flagship workflow, but the larger opportunity was a framework that could support many kinds of AI-assisted work. The same model extended across discovery, security, testing, and troubleshooting — different workflows, same mental model.

DISCOVERY

Prevent API sprawl before it starts

Search existing APIs and reusable assets before creating something new.

SECURITY

From reactive fixes to proactive guidance

Review APIs against organizational patterns and security practices before deployment.

TESTING

Generate coverage from API intent

Generate test cases, execute them, surface failures, and recommend fixes.

TROUBLESHOOTING

Investigate in chat. Resolve in context.

Diagnose issues conversationally, then return the actual fix to the API workspace.

09 — SCALING ACROSS PRODUCTS

One interaction model across the ecosystem

The first Agent experience launched in VS Code. I extended the interaction framework into API Connect and later API Studio. Each surface had different workflows and constraints, but the model remained consistent: start with intent, understand the plan, follow the work, act on the result in context.

API Manager

45%

API Studio

30%

VS Code

25%

What began as a solution for API creation became a reusable product pattern.

10 — OUTCOMES

Adoption showed where the model worked — and where it didn’t

78% of sessions ended in a successful outcome; 8% required escalation; 14% were abandoned. The abandonment rate mattered — rather than hiding it, I treated it as a roadmap signal for where longer or more complex workflows still created uncertainty.

TOP CAPABILITY

API Discovery

28% of all usage — reinforcing the research finding that developers wanted help understanding existing systems before creating something new.

COMMON WORKFLOW

Discover → Doc → Code

32% of multi-tool sessions followed this exact path — developers using the Agent to understand before they acted.

RETENTION

4.2×

Average repeat use within 30 days.

11 — WHAT I LEARNED

Trust is an interaction system

No single piece creates trust alone

A plan alone doesn’t create trust. A trace alone doesn’t create trust. A result alone doesn’t create trust. Together — predict before acting, see during execution, review before accepting — they do.

Patterns create more leverage than features

The broader impact came from establishing a framework that could scale across discovery, testing, security, troubleshooting, and future Agent experiences.

Constraints can improve the outcome

The limits of chat forced a clearer separation between explanation and action. That constraint produced a better model: the Agent explains the work, the product lets developers work with it.

The gap is the roadmap

The abandoned sessions became the clearest signal for what to design next.