This project is password-protected.
That’s not it — check your invitation.
Don’t have the password? Email me.
benner.dan@gmail.comCopy
A conversational agent for first responders, healthcare workers, caregivers, and teachers. Jobs like these take a mental toll, and typically don’t offer a way to recognize or deal with it.
This work was done inside a $100M stealth healthcare startup founded by former UnitedHealth Group leaders, where I drove the design thinking, guided the team’s direction, and influenced the product’s priorities.
The product does three things:
My work ran across the whole agentic system — who it was for, how it talked, how it was therapeutic, how it was built, and where it had to stop.
What follows is a mix: my thinking into the artifacts that drove the work, and the outcomes that shipped out of them.
Have a live conversation with the agent here, or use the full test bench to see all the decisions made by the agent.
01 — Origin and premise
Premise
Daily support for people whose jobs create accumulated stress: first responders, healthcare workers, government employees, caregivers.
Fifteen minutes a day of focused conversation, sustained over months, changes how people handle stress.
Business model
Unaddressed stress surfaces as burnout, lost quality of care, and turnover, and it lands on the employer’s . The bigger the issue, the bigger the cost, especially once chronic conditions take hold.
Clients pay nothing until their total cost of care drops, measured against employees who don’t use the product.
Therapeutic model
: future-oriented, goal-focused, evidence-based. I worked alongside the company’s consulting therapists to keep the conversation inside that discipline.
02 — Fit + Limits
Built for everyday cognitive load, not clinical conditions.
Mapping the gray areas the company hadn’t yet defined.
03 — The people
Reconstructed from real clients and the municipalities the product served, then used as seeds for synthetic personas and simulated conversations — the foundation for sample dialogues, scenario generation, and mapping how conversations actually go.
Jakob B.
Firefighter + Hazmat Technician
Missoula, MT
Primary pattern
Bravado as armor. Says “I'm good” even when he isn't.
Evelyn R.
DMV Document Examiner + Title Clerk
San Diego, CA
Primary pattern
Self-erasure. Cuts out her own needs first.
Kyler Z.
Advanced EMT
Cobb County, GA
Primary pattern
Deflection through busyness. Action instead of reflection.
04 — Personas Aren’t Enough
Every message a person types is dense with : emotional state, tone, narrative, motivation, hesitation, momentum. An LLM reads all of it at once, across thousands of dimensions, far more than any profile could hold.
Personas told us who someone is: demographics, scenario. They couldn’t tell us what state someone is in when they show up. Nielsen Norman’s archetype model (see right) gave me the principle that fixed that: read the user’s state, not just their identity, and adapt to it.
That principle scaled further than I expected. Personas describe people in categories humans can name; the model reads state, momentum, and pattern directly, including dimensions we have no words for. My job shifted from defining the user to describing to the model what to notice.
Impact: Everything below builds on this, from the scaffolding to the classifiers to the response model.
Nielsen Norman’s archetypes: a single read of user state. The product’s does the same thing at scale — reading the in every message across .
05 — Scaffolding
Three examples of how the product opens with different kinds of user. Two elements do the work: the Loading Screen Messages (shown on the loading screen, they set the tone before a word is exchanged) and the Opening Chat Prompts (HealthBot’s first move into conversation).
Basic Onboarding — Kyler
Short, low-pressure framing. No vocabulary the user has to look up.
Loading Screen Messages
Opening Chat Prompts
Light, low-commitment, no assumptions — he’s new.
Deep Engagement — Jakob
Belief-and-mindset framing for the user who is already past the surface.
Loading Screen Messages
Opening Chat Prompts
Reflective, invites depth without being corny — he’s engaged but guarded.
Retention — Evelyn
Identity-affirming language. The job here is to remind the user why they keep coming back.
Loading Screen Messages
Opening Chat Prompts
Warm, knowing, gently turns the lens back to her — she’s returning and puts herself last.
Three intros, same product, three different reads on what the user actually needs in the first 30 seconds.
05 — Scaffolding (Continued)
The model’s starts every topic light and only goes deeper when the user shows they can handle more. If they don’t, it stays light.
Impact: Self-awareness is the dimension I built out by hand (see right), so the same structure can be duplicated for any other dimension worth improving.
06A — Architecture
The words get used interchangeably. The difference decides everything about how this product behaves, because its users routinely say “I’m fine” when they’re not.
A respectful conversation is one that can aid self-discovery and self-reflection — which means the agent has to know what someone’s been working through, not start from zero every message.
That’s what I pushed for: a history people can look back on, goals they can return to. The original product had none of it. Every reply came from the last message alone, and leadership believed remembering a user’s circumstances would take away their agency. Without memory, this was an open journal disguised as a chatbot.
Impact: “Memory” became its own three-person team.
The chatbot matches a message. The agent reads a person.
See the Conversational AI Test Bench
06B — Architecture (Continued)
I routinely met with each ML engineer one on one. Everyone was working alone: services duplicated each other, handoffs broke daily, nobody owned the failures.
And the interviews surfaced the real finding — everyone was independently building toward the same unnamed thing, an agent. One engineer’s “conductor” was another’s “switch” or “mothership.”
Impact: The blueprints did two jobs. They showed us what we actually had: every service, what it did, who owned it.
And they gave us the 90-day cleanup plan, setting the stage for in-chat classifiers.
06C — Architecture (Continued)
The classifiers are a team, not a feature. Nine specialized ‘bots’, each with one job: reading atomic traits, tone, risk and safety, resilience, insights, narrative arcs, goals and motivations.
Each bot has a goal, a scope, and example insights it’s expected to produce, and they run in phases of increasing complexity, from granular signal detection up through narrative modeling.
07 — Product + Model Strategy
For an AI-first product, model design and product strategy are the same conversation. What the product is for and what the model is allowed to do collapse into one artifact — and if they don’t, one of them is fiction.
In and Out of Model’s Scope
| Category | Teach | Avoid |
|---|---|---|
| Therapeutic frameworks | basics | Clinical diagnosis |
| Emotions | Basic emotions, self-awareness | Telling users what they feel |
| Relationships | Self, work, interests | Prescribing what to do |
| Finance | Thinking about spending | Saving advice, financial planning |
| Safety | Crisis awareness, de-escalation | Handling the crisis itself |
| Religion | Only if user chooses | Guiding beliefs in any direction |
The pattern across every row: teach the skill, never perform it for them. The model builds capability, not dependency.
A more traditional product-level strategy, based on Google’s HEART framework
08 — Ideation + innovation
09 — Signal Detection (Phase I)
An ML notebook that turns open conversation into structured data the system can act on: how to pace the conversation, when to offer support, when to acknowledge growth. It reads a user’s language for behavioral signals — emotional cues, cognitive distortions, resilience markers, motivational patterns — each detected by its own narrowly scoped classifier.
The classifiers increase in complexity:
Impact: Each reports what it observes in the user’s language, with a score and a sample insight; the composite is what the decision layer acts on.
Kaggle Notebook (logic only) (login required)
10 — Response Generation (Phase II)
A notebook that takes Phase I’s output and shows how a reply gets built from what is underneath a message — emotional intensity, coping style, where the person is in their story — instead of its surface words.
Each response combines four parts:
Impact: A CBT-aligned reply built from behavioral evidence rather than keywords. Not diagnosis or advice, but the kind of thing a thoughtful friend with therapeutic training might say.
Kaggle Notebook (logic only) (login required)
11 — Human in the loop
The grading interface — thumbs up / thumbs down beside each response — collects human feedback on model outputs, the same approach major AI labs use to teach models what good responses look like.
Before it, model comparison was manual: the same prompts run against different models and scored by hand. Comparative model tests (XLSX)
Impact: The therapists are now the judges.
The model proposes; a person trained in the discipline decides what good looks like, one response at a time — and at scale, those judgments become the model’s behavior.
12 — Iteration
Examples of a different model behavior over a . Italics above each turn is what the model is recognizing or doing.
Evelyn — opening message
“I had a rough night. Mom didn't sleep so I didn't sleep. Now I'm at work and I can't focus.”
| Perceived | Decided | Suppressed | |
|---|---|---|---|
| Turn 1 | Caregiver fatigue, no margin | Acknowledge before extracting | Clinical-sounding questions (“rate your stress 1–10”) |
| Turn 2 | Self-criticism loop activating (“not doing anything well”) | Bring in prior session memory + ask for counter-evidence | “Have you considered respite care for your mother” — too solution-forward, user is in processing mode |
| Turn 3 | User supplied counter-evidence she didn't realize she had | Reflect it back and reframe depletion vs. failure | Praise that would feel hollow |
| Turn 4 | User accepted the reframe | Connect to a prior insight (guilt at day's end) and propose one small concrete action | A list of options (would feel like homework) |
13 — Working Demo
The test bench is a dual-pane view built for demoing the inputs for the entire agentic pipeline, not just chatting with it.
Demo the Test Bench
Live and interactive, running as a Node server against the Anthropic API — Demo the HealthBot Test Bench
14 — Future Intelligence
A future feature (unshipped): the user can opt into a framework they want to grow within — a mythology, a philosophy, a faith. The agent learns it and reflects through it, but never instructs. The user works things out inside a way of thinking they chose; the agent just knows it well enough to keep them company there.
It’s opt-in and set in settings, never inferred from conversation. Something this personal has to be chosen, not switched on quietly.
Kyler
Following the Hero’s Journey
Drawn to Star Wars or Marvel, Kyler opts into mythic-narrative reflection.
The agent never casts him as the hero or narrates his arc. It quietly frames things the way that structure does — the ordeal, the return — so a hard week reads as a chapter, not a conclusion.
Impact: Kyler recognizes the pattern himself; the agent just holds the shape.
The agency stays his.
Evelyn
Exploring a Wisdom Tradition
Curious about a tradition, or rooted in one, Evelyn opts into it.
The agent becomes fluent in that framework’s language of meaning, duty, and grace, and lets it inform how it reflects — offering the tradition’s questions, not its answers.
Impact: Whether Evelyn is new to the tradition or rooted in it, she is met in her own idiom and never preached to.
The agency stays hers.
Jakob
Focusing on a Philosophy or Practice
Stoicism, a coaching discipline, a way of thinking Jakob admires — he opts into it.
The agent reflects through that discipline’s habits (what’s in your control, what story you’re telling) without lecturing the doctrine.
Impact: The framework becomes the grain of Jakob’s conversation, not its content.
The agency stays his.