Research Automation

 

Research Corpus Automation: Keeping Knowledge Current at Scale

Project Timeline: 2026
Lead UX Designer | Systems Architecture and AI Governance
Presented at: AI Agentic Summit 2026
Link to presentation


Research deliverables are built for human readers: decks, journey maps, personas. To an AI, they're nearly unreadable. A bullet could be a finding, a header, or a joke. I designed and now operate a pipeline that treats structure as the deliverable and the polished version as a rendering of it, so findings stay current, traceable, and usable by both people and AI.


My Role

Identified the gap, designed the taxonomy and evidence-scoring methodology, and built the operating procedure that makes the process repeatable rather than dependent on one person's memory. Validated the approach on real research before formalizing it, and now run the pipeline as part of ongoing research operations.

Business Opportunity

We set out to address the following questions, and keep the answers forever up to date:

  • "What research has already been done on this topic?"

  • "Are these personas up to date?"

  • "What tools are our reps using?"

  • "How much of a problem is this pain point?"

“ This wasn’t about solving a UX research problem; but rather a knowledge currency problem.”
— Optimizing UX Deliverables for AI Consumption: Presentation to AI Agentic Summit, September 2026

Impact At-A-Glance

✓ Extraction, scoring, and comparison against the full corpus now run through Claude on real research
✓ Replaced subjective severity ratings with a three-factor evidence formula (Reach × Impact × Business Consequence), every score backed by a direct quote
✓ Reduced the system to two remaining manual touchpoints — triggering a session and applying the output — down from a fully manual, memory-dependent process

Example of a consolidated pain points view, rendered after processing hundreds of research artifacts from several different departments.

Process, Decisions, Solutions

Process

Raw research was the input. I started by imagining the presentation layer the output needed to become, then worked backwards to design a sustainable process connecting the two. The result is a five-step pipeline I built and now run: a researcher triggers a session with a new study; Claude extracts findings and scores them against the team taxonomy; Claude compares each finding against all five corpus assets — Personas, Pain Points Repository, Sales JTBD, Sales Segment Guide, and Tool Reference; Claude drafts the proposed diff and output; a person reviews and applies it.

Automation Process Overview

  1. Establish your foundational body of knowledge

  2. Feed new knowledge (in our case new research studies and findings) into the system, in our case an LLM.

  3. The system scans the existing assets that might contradict the new studies and offers recommendations

  4. Researcher approves recommendations and pushes the results to the live output or rendering

System Components

. 

Data Layer: Foundation of the Automation System

  • Scan targets/Living Items - checked against every finding and updated by it

  • Supporting systems - Track metadata and corroboration. These are the rules that govern our system.

  • Locked historical artifacts - this is the raw unstructured data that serve as snapshots in time

Tracing Research Through the System

Step 1 - Intake

A researcher uploads an new asset and the system logs it into the the inventory with with structured taxonomy: asset type, team, journey stage, participant role.

What actually gets uploaded today: raw transcripts (.vtt/.txt/.docx), research reports and readouts (.pdf/.docx/.pptx), spreadsheets and survey exports (.xlsx/.csv)


Step 2 - Extraction

The AI reads the transcript and systematically scores it using our pain point formula — Reach, Impact, Business Consequence — and tags it using only our controlled vocabulary.



Step 3 - Diff and Evidence Trail

The extracted findings get compared against our scan targets or our living foundational assets. It looks for where the new findings conflict or reinforce what is currently in our knowledge base. Items that don’t meet the evidence threshold get logged in the evidence tracker to build evidence for future changes.


Step 4 - Approve and Publish

Before anything goes live, the asset owner sees a before-and-after, with the actual source quote attached. Not "Claude changed this" — but why, and based on what. Approved updates immediately write to to the spreadsheet, and then downstream to the JSON which powers the React app and makes those changes viewable to stakeholders.


This is the payoff of "structured over visual" made concrete: everything the audience sees on the live site — persona pages, JTBD pages, topic pages — is a rendering. None of it is where the content actually lives.

Design Principles for AI Automation

The true value of our automation efforts. One upload to our system and several assets can be updated at once, keeping our knowledge base up to date.

Lessons Learned

Business Outcomes

✓ Turned a memory-dependent, ad hoc process into a governed pipeline that already runs on real research
✓ Corroborated independently by a separate large-scale research effort before most of the system was built
✓ Established a governance principle — AI handles the finding, humans handle the deciding — that's proving to generalize beyond UX research to any domain with a living source of truth and a human owner
✓ Current corpus: We’ve logged and extracted pain points from over 280 research reports or artifacts ensuring that these insights never get lost and could be retrieved again at a moment’s notice.