Enter password to view case study

Bridging Human & AI Knowledge

Bridging Human & AI Knowledge

Support teams often maintain two versions of the same knowledge — one for humans to read, one structured for an AI agent to draw from — updated on entirely separate schedules.

Our legacy setup kept these as two disconnected knowledge bases, with no mechanism linking a change in one to a change in the other. This meant an article could be corrected, retired, or superseded by a human agent while the AI-facing version stayed untouched for weeks, causing the AI agent to give customers outdated answers or escalate questions a human had already resolved, with no reliable way to see where the two knowledge bases had drifted apart.

ROLE

Lead Product Designer

Lead Product Designer

COMPANY

ANNA Money

ANNA Money

Timeline

2026

2026

SKILLS

Product Design, Workflow Design, AI Assisted Support, Systems Thinking, Knowledge Systems

Product Design, Workflow Design, AI Assisted Support, Systems Thinking, Knowledge Systems

5x growth

In AI Knowledge

82% converted

from Human knowledge → AI

Decrease in escalations

AI Agent conversations increasing

The challenge

The company ran two knowledge bases side by side: a human-readable one that support agents used every day, and a structured one that fed an AI customer agent.

They started life as separate things, maintained on separate schedules — which meant the AI agent was quietly answering customers from information that support agents had already corrected, retired, or superseded weeks earlier.

Every gap was a customer hallucinated answers, or the AI agent escalating a question a human, even though the answer existed already answered in the knowledge base.

The fix wasn't "write more AI content."

It was closing the loop — making sure that every time a human article changed, the AI-facing version changed with it, without turning every content update into a manual double-entry job for already-stretched support writers.

My role

I owned three connected pieces of this:

  • The workflow agents follow to keep the AI knowledge base current
    What triggers a review, who approves it, what “done” looks like, and how agents actually make the resulting content updates.

  • Shaping the translation logic
    The rules that decide what counts as customer-relevant, and how a human-facing article gets rewritten for an AI reader.

  • The knowledge gaps dashboard
    A regularly-refreshed view showing where the AI agent was actually failing customers, where key gaps were for each business area and whether work was closing those gaps.

Making the agent workflows more user friendly

The first version of the workflow used Cursor and GitHub to make knowledge base updates. It gave agents a powerful way to work directly with the underlying content, but it also introduced a new problem: the interface was technical.

For agents who weren't comfortable working in a code environment, the prospect of changing files and potentially breaking something was daunting. Once they became familiar with GitHub, the workflow became much easier, but there was still a learning curve before they could use it confidently.

We evolved this into a Claude skill in Slack, allowing agents to make the same kinds of updates through natural-language prompts. Instead of navigating a code environment, they could simply describe what they wanted to create or change.

This reduced the technical barrier between identifying a knowledge gap and actually fixing it — making the workflow much closer to the way agents already worked.

That last part was particularly important. If reviewers stop trusting a “nothing to do here” verdict, they re-read everything manually anyway. The workflow then adds overhead instead of removing it.

*Illustrative data for portfolio use — figures are representative, not actual production numbers.

Designing the translation workflow

The two knowledge bases were structured completely differently, the AI version didn't map 1:1 to article paths in the human one, so "just diff the two folders" was never going to work.

I designed the process around a simple principle: every human knowledge base change gets a proposed AI counterpart, but nothing ships without a human seeing the diff first.

  • The Trigger - what changes should initiate the process.

  • The proposal - its own linked review, never silently merged.

  • The approval gate - a person reviews the actual before/after content, not a checkbox.

  • The “do nothing” edge case - because not every human edit is customer-relevant.

That last part was particularly important. If reviewers stop trusting a “nothing to do here” verdict, they re-read everything manually anyway. The workflow then adds overhead instead of removing it.

*Illustrative data for portfolio use — figures are representative, not actual production numbers.

Shaping the translation logic

A straight copy-paste didn't work. The human knowledge base was written by agents, for agents, so the content contained context and language that made sense to a human reader but wasn't appropriate for an AI agent.

The rules I helped define included:

  • Strip everything internal — screenshots, internal tool references, “raise a task to…”, and other agent-only instructions.

  • Rewrite the point of view — human articles might say “you can find this in Settings”, but the AI knowledge base isn't read by the customer. It’s read by the AI agent, which then talks to the customer. So the content becomes “the customer can find this in Settings”.

  • Respect explicit “internal-only” markers — content marked as internal should never be translated into the customer-facing AI knowledge base.

  • Structure content for retrieval — a heading that reads as a real customer question retrieves far better than a generic one.

The goal wasn't to reproduce the human knowledge base in another format. It was to translate the knowledge into something the AI could retrieve and use correctly.

Understanding the gaps with dashboards

Working with a backend engineer on the data pipeline — primarily from real customer conversations and knowledge gaps — I designed a dashboard that turned raw escalation logs into a ranked, regularly-updated view of where the AI was struggling.

It showed which topics were driving the most human hand-offs, how that trended over time, and whether the knowledge base updates we shipped actually moved the number.

Each topic could then be expanded into the specific knowledge gap, what the AI-facing content currently said, work in progress against it, and a link into deeper analysis.

Agents could also use the dashboard to implement a gold-answer update using a single command string, allowing them to address an identified AI knowledge gap without having to create or write a new article from scratch.

*Illustrative data for portfolio use — figures are representative, not actual production numbers.

Results

The AI-facing knowledge base grew substantially faster than the human knowledge base, both in overall word volume and article content.

More importantly, the gap between the two knowledge bases narrowed measurably, translating into fewer customers receiving outdated answers from the AI agent.

What I'd do differently

Build trust in the “no change needed” verdict earlier.
Getting that verdict trustworthy enough that reviewers stopped double-checking took a couple of iterations. I’d invest in that earlier next time, because the value of the automation depends on people trusting it enough to stop doing the work manually.

Connect the dashboard to the propagation workflow from day one.
I’d wire the dashboard directly into the propagation workflow, rather than leaving “we closed this gap” and “the dashboard reflects it” as two separate steps. Closing the loop earlier would make the system more coherent and reduce unnecessary handoffs.

Revisit the approval process as the system evolves.
Over time, more signals were introduced, which meant more reviews and more notifications were added to the process. What started as useful oversight became increasingly overwhelming. I’d review the approval model as the system matured, looking at which signals genuinely need human attention and which can be safely consolidated or automated.