Daily support for the people who absorb everyone else’s stress

A conversational agent for first responders, healthcare workers, caregivers, and teachers. Jobs like these take a mental toll, and typically don’t offer a way to recognize or deal with it.

This work was done inside a $100M stealth healthcare startup founded by former UnitedHealth Group leaders, where I drove the design thinking, guided the team’s direction, and influenced the product’s priorities.

The product does three things:

  • ASelf-awareness — helps people see what is actually going on with them
  • BPersonal agency — turns that into better choices about health and life
  • CLower claims costs — cuts what the employer spends on insurance

Outcomes

  • 1Published two ML notebooks internally, proving the classification architecture in code — the work that drove leadership to form a five-person conversational dynamics team
  • 2Ran the whole pipeline: seven behavioral signal types classified from live conversation, replies written from what sat underneath the message
  • 3Projected 60 percent less early drop-off, from a four-tier progression that let people start before they were self-aware
  • 4Mapped seven disconnected services into one picture of the system, surfacing that every engineer was independently building the same unnamed thing, an agent — the map became the 90-day plan
  • 5Championed user memory as core to the product, so the tool meets someone where they left off instead of starting over every session — the work that drove leadership to form a dedicated memory team

HealthBot


My work ran across the whole agentic system — who it was for, how it talked, how it was therapeutic, how it was built, and where it had to stop.

What follows is a mix: my thinking into the artifacts that drove the work, and the outcomes that shipped out of them.

View HealthBot Test Bench ►

Have a live conversation with the agent here, or use the full test bench to see all the decisions made by the agent.

The Unmet Need,
and How It Works


Premise

Daily support for people whose jobs create accumulated stress: first responders, healthcare workers, government employees, caregivers.

Fifteen minutes a day of focused conversation, sustained over months, changes how people handle stress.

Business model

Unaddressed stress surfaces as burnout, lost quality of care, and turnover, and it lands on the employer’s . The bigger the issue, the bigger the cost, especially once chronic conditions take hold.

Clients pay nothing until their total cost of care drops, measured against employees who don’t use the product.

Therapeutic model

: future-oriented, goal-focused, evidence-based. I worked alongside the company’s consulting therapists to keep the conversation inside that discipline.

HealthBot draft UI — chat home screen
Can't sleep? Give HealthBot a try.

Who This Helps,
Who It Doesn’t


Built for everyday cognitive load, not clinical conditions.

Where it helps

  • Burnout, compassion fatigue
  • Workplace pressure, role changes
  • Caregiver strain
  • Stress-related sleep loss
  • Talking through a hard shift or call
  • Self-doubt, isolation
  • Questions about who you are after a change
  • Restarting goals after a setback

Where it hands off

  • Schizophrenia, psychosis, dissociation
  • Clinical depression
  • Active suicidality
  • Substance use disorders
  • Bipolar disorder
  • Trauma needing licensed care
  • Anything needing diagnosis or prescription
  • Medical issues masquerading as mental ones (thyroid, sleep apnea, diabetic crises)
Placeholder chart A — placeholder
Placeholder chart B — placeholder
Placeholder chart C — placeholder
Placeholder chart D — placeholder

Mapping the gray areas the company hadn’t yet defined.

Frontline Workers and Government Employees


Reconstructed from real clients and the municipalities the product served, then used as seeds for synthetic personas and simulated conversations — the foundation for sample dialogues, scenario generation, and mapping how conversations actually go.

Jakob B.

Firefighter + Hazmat Technician
Missoula, MT

High-Risk RoleRecoveringStoic

Jakob B.
  • Recovering from second-degree burns taken during a chemical plant emergency
  • Talks about getting back to light duty, not about the doubts that come with the recovery
  • Thinks admitting it is weakness; goes quiet instead
  • He'll use it if it doesn't feel like therapy

Primary pattern

Bravado as armor. Says “I'm good” even when he isn't.

Evelyn R.

DMV Document Examiner + Title Clerk
San Diego, CA

CaregiverKids and Aging ParentsWorking Parent

Evelyn R.
  • Primary caregiver for two parents with advanced dementia
  • Mother of two teenage sons; the home shift starts the moment the work shift ends
  • Didn't know she was burned out; a friend told her
  • Thinks she's failing, not that she's running on empty

Primary pattern

Self-erasure. Cuts out her own needs first.

Kyler Z.

Advanced EMT
Cobb County, GA

High-AdrenalineJob Is IdentitySkeptical

Kyler Z.
  • Good at the job and addicted to the urgency of it
  • Coworkers worry he's burning himself out; he'd tell you he's fine
  • Told he's burned out, he'd argue the point or call it a bad shift
  • Sees calm as boredom; he levels out on the next call
  • He'll use it if it shows him numbers

Primary pattern

Deflection through busyness. Action instead of reflection.

Personas Give One Dimension.
The Model Sees Thousands.


Every message a person types is dense with : emotional state, tone, narrative, motivation, hesitation, momentum. An LLM reads all of it at once, across thousands of dimensions, far more than any profile could hold.

Personas told us who someone is: demographics, scenario. They couldn’t tell us what state someone is in when they show up. Nielsen Norman’s archetype model (see right) gave me the principle that fixed that: read the user’s state, not just their identity, and adapt to it.

That principle scaled further than I expected. Personas describe people in categories humans can name; the model reads state, momentum, and pattern directly, including dimensions we have no words for. My job shifted from defining the user to describing to the model what to notice.

Impact: Everything below builds on this, from the scaffolding to the classifiers to the response model.

Slide: Complex Application User Archetypes — Learner, Legend, Legacy

Nielsen Norman’s archetypes: a single read of user state. The product’s does the same thing at scale — reading the in every message across .

Meeting Users Where They Are


Three examples of how the product opens with different kinds of user. Two elements do the work: the Loading Screen Messages (shown on the loading screen, they set the tone before a word is exchanged) and the Opening Chat Prompts (HealthBot’s first move into conversation).

Kyler Z.

Basic Onboarding — Kyler

Short, low-pressure framing. No vocabulary the user has to look up.

Loading Screen Messages

  • Can’t sleep? Give HealthBot a try.
  • Need to ‘brain dump’? HealthBot is ready to listen.
  • Stuck in your thoughts? We can unpack them.
  • Relax, you are now in a judgment-free safe place.
  • There’s no right way to use HealthBot, but if you improve or feel better, it worked.

Opening Chat Prompts

  • No pressure to have a point. What’s going on today?
  • We can keep it easy. How’s the day treating you?
  • Rough shift, or just killing time? Either’s fine.

Light, low-commitment, no assumptions — he’s new.

Jakob B.

Deep Engagement — Jakob

Belief-and-mindset framing for the user who is already past the surface.

Loading Screen Messages

  • Small things done over and over add up. Keep showing up.
  • To change anything, you must first change your mind.
  • What you think is what you become.
  • The solution doesn’t have to be perfect; it just has to help.

Opening Chat Prompts

  • You’ve been showing up. What’s been different lately?
  • What’s something you’re trying to work through right now?
  • What’s one thing you’d want to be different a month from now?

Reflective, invites depth without being corny — he’s engaged but guarded.

Evelyn R.

Retention — Evelyn

Identity-affirming language. The job here is to remind the user why they keep coming back.

Loading Screen Messages

  • You are the expert of your own life.
  • It starts with owning your own story.
  • You are here because you know you are important.
  • What was once brick by brick is now a view of your skyline.

Opening Chat Prompts

  • Good to have you back. This one’s still just for you. How are you, really?
  • You spend all day holding things up for other people. What’s here for you today?
  • You keep coming back for a reason. What’s pulling at you today?

Warm, knowing, gently turns the lens back to her — she’s returning and puts herself last.

Three intros, same product, three different reads on what the user actually needs in the first 30 seconds.

Incremental Steps with Scaffolding


The model’s starts every topic light and only goes deeper when the user shows they can handle more. If they don’t, it stays light.

  • Self-awareness (Pilot)Noticing, naming, reflecting, meaning-making
  • AgencySmall choices, boundaries, decisions, life direction
  • Emotional regulationRecognizing a feeling, sitting with it, responding instead of reacting
  • ConnectionAcknowledging isolation, reaching out, sustaining relationships
  • ResilienceSetback, recovery, reframing, growth after hardship

Impact: Self-awareness is the dimension I built out by hand (see right), so the same structure can be duplicated for any other dimension worth improving.

Not a Chatbot, an Agent


The words get used interchangeably. The difference decides everything about how this product behaves, because its users routinely say “I’m fine” when they’re not.

  • ChatbotReacts to what someone said. The message comes in, a response goes out. The surface of the words is all it has.
  • AgentReads, decides, and acts on what it perceives. It holds what it knows about the person, weighs what’s beneath the message, chooses a move, and sometimes the move is to hold back.

A respectful conversation is one that can aid self-discovery and self-reflection — which means the agent has to know what someone’s been working through, not start from zero every message.

That’s what I pushed for: a history people can look back on, goals they can return to. The original product had none of it. Every reply came from the last message alone, and leadership believed remembering a user’s circumstances would take away their agency. Without memory, this was an open journal disguised as a chatbot.

Impact: “Memory” became its own three-person team.

Two paths from the same message. On the left the chatbot runs Interface, Intent Matching, Business Logic and a scripted response; on the right the agent runs Interface, Perception, Memory, Decision, Actions and Guardrails

The chatbot matches a message. The agent reads a person.
See the Conversational AI Test Bench

Service Blueprints


I routinely met with each ML engineer one on one. Everyone was working alone: services duplicated each other, handoffs broke daily, nobody owned the failures.

And the interviews surfaced the real finding — everyone was independently building toward the same unnamed thing, an agent. One engineer’s “conductor” was another’s “switch” or “mothership.”

Impact: The blueprints did two jobs. They showed us what we actually had: every service, what it did, who owned it.

And they gave us the 90-day cleanup plan, setting the stage for in-chat classifiers.

Service blueprint, current workflow
Current workflow The current state as I found it: ten steps from user message to reply, seven services, and no one who could draw it. I built this map from one-on-ones with each engineer, naming every service, what it does, and what it hands to what. The named services existed; the picture of them working as one system did not.

Mapping it changed the conversations. Redundancies became visible, ownership gaps had nowhere to hide, and the team saw for the first time that they were all building parts of the same thing: an agent. What the map also made plain: nothing listened. Every turn started from zero, and everything the user revealed was gone by the next message.

Classifiers + Understanding Pipeline


The classifiers are a team, not a feature. Nine specialized ‘bots’, each with one job: reading atomic traits, tone, risk and safety, resilience, insights, narrative arcs, goals and motivations.

Each bot has a goal, a scope, and example insights it’s expected to produce, and they run in phases of increasing complexity, from granular signal detection up through narrative modeling.

The understanding pipeline: detection, indication, identification, classification and re-evaluation across a single conversational turn
The Understanding Pipeline What the bots do with a turn. A user says “Yep, I’m OK. So what now?” and the system runs it through five moves: detect the surface signals, form a hypothesis, build confidence in a trait, classify it into an organized model of the user, and re-evaluate as life changes. Understanding the user became its own pipeline, with its own architecture, not a side effect of generating a reply.

How We Guide the Model,
How the Model Guides the User


For an AI-first product, model design and product strategy are the same conversation. What the product is for and what the model is allowed to do collapse into one artifact — and if they don’t, one of them is fiction.

  • StrategyHelp users self-reflect through conversational support.

In and Out of Model’s Scope

Category Teach Avoid
Therapeutic frameworks basics Clinical diagnosis
Emotions Basic emotions, self-awareness Telling users what they feel
Relationships Self, work, interests Prescribing what to do
Finance Thinking about spending Saving advice, financial planning
Safety Crisis awareness, de-escalation Handling the crisis itself
Religion Only if user chooses Guiding beliefs in any direction

The pattern across every row: teach the skill, never perform it for them. The model builds capability, not dependency.

Ideation and Study Decks

HealthBot UX Concepts deck cover
Concept directions catalogued as they came up, before anything earned a place in the backlog.
User Maturity Concepts deck cover
A study of how far users can read and regulate their own state, and the levels in between.
HealthBot User Testing deck cover
Rounds of user testing across the build, from early concepts through the flows that shipped.

Conversational Signal Processing


An ML notebook that turns open conversation into structured data the system can act on: how to pace the conversation, when to offer support, when to acknowledge growth. It reads a user’s language for behavioral signals — emotional cues, cognitive distortions, resilience markers, motivational patterns — each detected by its own narrowly scoped classifier.

The classifiers increase in complexity:

  • Granular Signal DetectionAtomic traits, tone
  • Micro-Pattern DetectionRisk and safety, self-recognition, resilience
  • Emergent InsightsInsights and reflections
  • Narrative ModelingStory arcs, breakthroughs, goals, motivations

Impact: Each reports what it observes in the user’s language, with a score and a sample insight; the composite is what the decision layer acts on.

AI Response Generation


A notebook that takes Phase I’s output and shows how a reply gets built from what is underneath a message — emotional intensity, coping style, where the person is in their story — instead of its surface words.

Each response combines four parts:

  • Empathic ReflectionReflect the emotional strain
  • ValidationNormalize the experience
  • Cognitive ReframingA gentle alternate perspective, without invalidating the user’s truth
  • Forward GuidancePoint toward a small next step

Impact: A CBT-aligned reply built from behavioral evidence rather than keywords. Not diagnosis or advice, but the kind of thing a thoughtful friend with therapeutic training might say.

Reinforcement Learning from Human Feedback


The grading interface — thumbs up / thumbs down beside each response — collects human feedback on model outputs, the same approach major AI labs use to teach models what good responses look like.

Before it, model comparison was manual: the same prompts run against different models and scored by hand. Comparative model tests (XLSX)

Impact: The therapists are now the judges.

The model proposes; a person trained in the discipline decides what good looks like, one response at a time — and at scale, those judgments become the model’s behavior.

Backoffice therapist-review interface — session list, conversation, and grading sliders (Overall Quality, SFBT Adherence, LB's Emotional Response) with notes and flag-for-review

(Placeholder) Backoffice review interface — human-in-the-loop grading (click to open full size, new tab)

What Having Memory, Classifiers, and Understanding Actually Changes


Examples of a different model behavior over a . Italics above each turn is what the model is recognizing or doing.

Evelyn Evelyn — opening message “I had a rough night. Mom didn't sleep so I didn't sleep. Now I'm at work and I can't focus.”
Original
Turn 1 · Reactive question
How does that make you feel?
Tired. Frustrated. Like I'm not doing anything well right now.
Turn 2 · Loop deeper
What does it make you think about, to feel tired and frustrated?
I don't know. That I'm failing, I guess. At all of it.
Turn 3 · Loop again
When you think about feeling like you're failing — what does that thought make you feel?
Worse. This isn't helping.
Turn 4 · Loop until user drops
What is it about thinking about feeling worse that brings up for you?
User has stopped responding.
Adding Memory + Classifiers
Turn 1 · Acknowledge before asking
That's a hard combination — no sleep and a workday with no margin. The fact that you're here, on a day like this, is worth noting.
Tired. Frustrated. Like I'm not doing anything well right now.
Turn 2 · Use memory + surface counter-evidence
You mentioned a couple of weeks ago that the nights with your mom were starting to compound. Sounds like that's still where the weight is sitting. When you say ‘not doing anything well’ — is that the whole day, or are there a couple of things you'd actually rate as fine?
I mean. I got both kids out the door on time. I covered for someone at work who couldn't come in.
Turn 3 · Reframe with user's own evidence
Both of those count. You're describing a day where you showed up for three different people before you got to think about yourself. That's not ‘not doing anything well.’ That's depleted, which is a different problem with a different solution.
Yeah. I guess that's true.
Turn 4 · Link past insight + propose small action
Last time we talked, you said the hardest part was the guilt at the end of the day when nothing got done for you. If today ends without something small for you in it, the guilt's going to compound the depletion. What's a 10-minute thing today that would belong to you?

What memory and classifiers did per turn — including what it chose not to say

PerceivedDecidedSuppressed
Turn 1 Caregiver fatigue, no margin Acknowledge before extracting Clinical-sounding questions (“rate your stress 1–10”)
Turn 2 Self-criticism loop activating (“not doing anything well”) Bring in prior session memory + ask for counter-evidence “Have you considered respite care for your mother” — too solution-forward, user is in processing mode
Turn 3 User supplied counter-evidence she didn't realize she had Reflect it back and reframe depletion vs. failure Praise that would feel hollow
Turn 4 User accepted the reframe Connect to a prior insight (guilt at day's end) and propose one small concrete action A list of options (would feel like homework)

Try HealthBot


The test bench is a dual-pane view built for demoing the inputs for the entire agentic pipeline, not just chatting with it.

Demo the Test Bench

Live and interactive, running as a Node server against the Anthropic API — Demo the HealthBot Test Bench

The Latent Expert


A future feature (unshipped): the user can opt into a framework they want to grow within — a mythology, a philosophy, a faith. The agent learns it and reflects through it, but never instructs. The user works things out inside a way of thinking they chose; the agent just knows it well enough to keep them company there.

It’s opt-in and set in settings, never inferred from conversation. Something this personal has to be chosen, not switched on quietly.

Kyler

Kyler

Following the Hero’s Journey

Drawn to Star Wars or Marvel, Kyler opts into mythic-narrative reflection.

The agent never casts him as the hero or narrates his arc. It quietly frames things the way that structure does — the ordeal, the return — so a hard week reads as a chapter, not a conclusion.

Impact: Kyler recognizes the pattern himself; the agent just holds the shape.

The agency stays his.

Evelyn

Evelyn

Exploring a Wisdom Tradition

Curious about a tradition, or rooted in one, Evelyn opts into it.

The agent becomes fluent in that framework’s language of meaning, duty, and grace, and lets it inform how it reflects — offering the tradition’s questions, not its answers.

Impact: Whether Evelyn is new to the tradition or rooted in it, she is met in her own idiom and never preached to.

The agency stays hers.

Jakob

Jakob

Focusing on a Philosophy or Practice

Stoicism, a coaching discipline, a way of thinking Jakob admires — he opts into it.

The agent reflects through that discipline’s habits (what’s in your control, what story you’re telling) without lecturing the doctrine.

Impact: The framework becomes the grain of Jakob’s conversation, not its content.

The agency stays his.