Nextiva · Healthcare Contact Center

Case study contents

    UX Case Study · Enterprise AI · Healthcare

    From fragmented contact-center workflows to one trusted AI workspace.

    Designing an AI scheduling copilot that helps healthcare agents resolve requests faster without taking control away.

    nextiva · agent workspace live call, mid-booking
    The unified agent workspace mid-call, with NextIQ drafting a suggested action and pre-filling the booking
    Actual product screen: the live call transcript, the copilot's suggested action, and the pre-filled booking — side by side.
    Company
    Nextiva
    Role
    Lead Product Designer
    Product
    Healthcare Contact Center — NextIQ
    Pilot scale
    3 practices · ~40 agents
    Stage
    Discovery → MVP pilot
    01

    Executive summary

    Healthcare agents were drowning in software instead of listening to patients.

    I consolidated four fragmented systems into one canvas. When a moderated usability study showed that wasn't enough, I evolved it into an AI copilot NextIQ that listens to the live call and does the searching, so the agent can focus on the patient.

    But this is healthcare, so the copilot never became the only way through. The manual scheduler stayed in place as the safety floor — every AI action pre-fills it, and any agent can take manual control at any time. It shipped as an MVP pilot, proven with a moderated study plus impact projections modeled on published benchmarks, built on one non-negotiable: the AI recommends, the human decides.

    Measured in testing (n=10):

    SUS score
    61 → 86
    Unaided completion
    4/10 → 10/10
    Time-on-task
    4m20s → 1m30s
    Holds during task
    7/10 → 0
    Cognitive load (TLX)
    74 → 39
    02

    My role

    Lead Product Designer — the single design owner end-to-end, and design lead across a cross-functional pod of Product, Engineering, and a two-person UX Research team.

    Strategy & ownership

    Led design for scheduling, prescriptions, and bill pay; translated research into a direction and defended it, including a pivotal disagreement over how much to automate (Section 14).

    Craft

    Flows, wireframes, hi-fi screens, and the AI copilot interaction model, authored from first principles.

    Cross-functional leadership

    Daily feasibility alignment with PM and Engineering; partnered with Research to keep every decision evidence-based.

    03

    The business problem

    Voice is still the dominant channel in healthcare. Every minute an agent spends hunting across systems is a longer hold, a longer call, and a higher chance the patient calls back or leaves.

    Patient identity, scheduling, prescriptions, and billing each lived in a different system including EHRs we didn't control. To finish one request on a live call, the agent became the search engine: switching across five apps and screens on average, using the hold button as a crutch, and holding the entire workflow in their head under time pressure, where being wrong isn't an option.

    The problem wasn't a badly designed screen. It was that the work had no home.

    A front-desk healthcare agent, head in hands, at a desk with two monitors covered in sticky notes

    Market context (Nextiva's published healthcare research not results from this project): ~60% of calls to smaller practices go unanswered; 76% of healthcare leaders report being overwhelmed by administrative workload; 1 in 3 patients leave a provider after a single bad experience.

    04

    Research & discovery

    A two-person UX Research team ran a mixed-methods study user interviews, contextual inquiry, workflow analysis, and moderated usability testing. I wasn't a passenger: I helped shape the questions, sat in on every session, pressure-tested findings against feasibility, and owned the translation from insight to design decision.

    I started inside the system agents actually lived in.

    Before I designed a single scheduler screen, I studied the incumbent EHR. I collected and annotated the real scheduling, appointment, and patient-workflow screens the fields, the search logic, the status states and mapped how a booking truly moved through the system, branch by branch.

    Annotated study — the incumbent EHR, screen by screen
    appointments schedule daily grid
    Incumbent EHR appointments schedule grid, showing patients, visit status, room, and visit type across multiple clinics
    The daily schedule grid — every visit, its status, room, and type, across multiple clinic locations.
    appointment schedule create / edit dialog
    Incumbent EHR appointment schedule dialog with date, provider, location, appointment type, duration, and reason fields
    The booking dialog itself — provider, location, type, duration, reason. The field order that shaped my scheduler.
    provider search — find an open slot
    Incumbent EHR search dialog for finding the earliest open slot for a provider and location
    Finding an open slot doctor, location, day-of-week, and time-range filters, searched one provider at a time.
    patient workflow — status states
    Incumbent EHR patient workflow status dropdown, showing states like scheduled, running late, arrived, no show
    The status vocabulary agents worked in scheduled, running late, arrived, no-show the states my system needed to respect.

    That study is why the scheduler I designed fit clinical reality instead of an idealized version. The specialization → provider → visit type → location → slot order and the verification gate both came directly from how the incumbent system structured the work — and grounding the redesign in the system agents already knew made the consolidation argument concrete to stakeholders.

    flow mapping — scheduling journey, per intake path
    Mapping the existing scheduling workflow across visit type, location preference, date/time, and provider selection
    The workflow I was replacing — one "book an appointment" request touching identity, insurance, availability, and location, each in a different place. This map became my argument for consolidation.

    What we learned:

    • Constant system-switching agents touched an average of five apps to finish one request.
    • Fragmented patient info the full picture never lived in one place.
    • The hold button as a crutch agents parked patients just to go find things.
    • Call-backs as failure many requests couldn't close in one call, and most patients who don't reach a person on the first try never call back.
    • They wanted one workspace not another tab to alt-tab into.
    05

    Goals

    One canvas

    End the alt-tab between scheduling, prescriptions, billing, and context.

    Lower cognitive load

    The system carries the memory work; the agent carries the conversation.

    Finish in one call

    Fewer holds, fewer call-backs.

    Trustworthy AI

    Explainable, correctable, always subordinate to the human.

    Safe by default

    HIPAA and clinical reality at every step with a manual path that always works.

    The framing question

    How might we let an agent complete a request in a single, unbroken conversation without holding the entire system in their head, and without ever depending on the AI being right?

    06

    What I built first the manual scheduler

    The roadmap was scheduling-first. I consolidated the fragmented tools into one workspace and gave scheduling a single structured flow specialization → provider → consultation type → location → date → slot with rescheduling reusing the pattern and the current appointment pinned so the agent never lost the thread.

    One canvas, no new tab

    The scheduler opened inside the call workspace, not as another system to alt-tab into.

    Pinned context on reschedule

    The current appointment stays visible, so the agent isn't holding it in memory mid-call.

    Structured over free-form

    A guided path, modeled on the incumbent EHR's real sequence, to cut booking errors where a wrong slot has clinical consequences.

    The manual scheduler — a real booking, start to finish
    01 · open the schedule appointment modal
    Schedule appointment modal, empty, with specialization, appointment type, doctor preference, consultation type, and location fields
    01 · Specialization, appointment type, doctor, consultation type, location — five decisions before a single date is picked.
    02 · agent picks a preferred date, slots load
    Schedule appointment modal with a date range picked and available slots loading
    02 · Fields complete, a preferred day selected — the system goes looking for open slots.
    03 · available slots returned
    Schedule appointment modal showing a grid of available time slots for the selected day
    03 · A full day of slots, one tap from booked — once the agent has read them all out loud.
    04 · patient asks for a different week
    Schedule appointment modal re-querying availability after the patient asks for a different week
    04 · The patient changes their mind the agent re-queries, and the wait starts over.
    05 · results shown for the new week
    Schedule appointment modal showing a new set of results after re-querying availability
    05 · New slots, same fifteen-field form every revision costs another lap through the fields.
    06 · no availability for the requested day
    Schedule appointment modal showing no appointments available for the selected day, with a next-available link
    06 · No slots that day — the dead end every agent eventually hit, patient still on the line.

    It killed the alt-tabbing. It also exposed the ceiling of "organize the chaos better."

    07

    Patient verification the HIPAA hard gate

    Before any request is actioned, the agent has to prove they're talking to the right patient. In a system that would eventually let an AI pre-fill actions, identity verification is the one place I designed to be un-skippable.

    • Inline, glanceable check for returning callers — DOB · MRN · ZIP, one field at a time, without stopping the conversation.
    • Full blocking modal for new patients, or the moment any field mismatches the hard stop.
    Identify before the agent even speaks
    live mode — inbound call transferred from the IVA
    Live mode active, inbound call pop transferred from the virtual assistant with caller name and number
    The call arrives already matched — transferred from the virtual assistant with caller identity in motion.
    phone number matches multiple patients on file
    Patient verification modal showing the incoming number belongs to multiple existing members, with a list to select who identified themselves
    When one phone number covers a household, the agent confirms which patient is actually on the line before anything else happens.
    Verify — one field at a time, until the gate opens
    verification requirements nothing confirmed yet
    Patient verification modal listing date of birth, MRN ID, and ZIP code, none yet checked, Continue disabled
    DOB, MRN, ZIP — required, and nothing pre-checked. Continue stays disabled until they are.
    one field confirmed — DOB matched against the caller
    Patient verification modal with date of birth checked off after matching the caller, MRN and ZIP still pending
    Date of birth checked off as the agent hears it match — the glanceable, one-field-at-a-time check.
    all three fields verified — gate opens
    Patient verification modal with all three fields checked and Continue enabled
    All three verified, Continue unlocked. The AI is never allowed to act until identity matches.

    I walked this condensed flow through compliance stakeholders before pilot. They signed off because the hard gate — mismatch → full modal → frozen pre-fill — stayed intact.

    08

    The test that forced the pivot

    Tidy forms are still forms. The agent was still the search engine reading availability, typing while the patient waited. So we tested the manual scheduler properly, and the numbers were blunt.

    Manual scheduler moderated usability study (n = 10 agents):

    • SUS: 61 — below the 68 benchmark; "OK, not good."
    • Unaided completion: 4 of 10 agents booked a follow-up without moderator help.
    • Median time-on-task: 4m 20s.
    • Holds: 7 of 10 sessions still put the (simulated) patient on hold.
    • Verbatim: "not usable as is" for scheduling and refills.

    Illustrative quotes below representative of session feedback, not verbatim transcripts.

    Time on task

    "You've made it neater. I'm still doing all the thinking out loud while they sit on hold."

    ★★★☆☆

    Agent 03

    Front-desk, Practice B

    Missing calendar

    "I still have to hold their whole schedule in my head to know if a slot even makes sense."

    ★★★☆☆

    Agent 07

    Front-desk, Practice A

    When every participant hits the same structural wall, that's not a preference it's a signal. It told me the next move wasn't a better form. It was a different kind of help.

    research readout key findings
    Key findings slide: participants state product is not usable as is for scheduling or refilling, missing calendar, unrealistic prescription flow
    The verbatim readout — "not usable as is," plus the gaps: missing calendar, unrealistic refill flow, thin verification.
    research readout — missing information
    Missing information slide detailing gaps in prescription modal, calendar, and verification process
    The gap analysis — raw material for both the fix and the V2 backlog.
    09

    The pivot — and why the manual scheduler stayed

    Don't make the search faster do the search for the agent.

    The study pointed at two moves, not one. First, I closed the structural gaps it exposed — a real calendar, a realistic refill flow, tighter verification. Second, I layered a copilot on top to remove the searching entirely. NextIQ listens to the live transcript, identifies intent, and surfaces the appointment, prescription, or billing action the agent would otherwise go dig for.

    Here's the decision that matters most in a healthcare product: the improved manual scheduler was not thrown away. It became the floor the AI sits on.

    Every copilot suggestion pre-fills that same scheduler, and the agent can always take the wheel — edit any field, or ignore the AI entirely and book by hand.

    The copilot never took the wheel. In healthcare you don't want a single point of failure that happens to be an AI — the human must always be able to complete the task without it. That principle became the line I fought to hold (Section 14).

    Manual onlyFinal · AI-assisted
    Finding availabilityAgent searches by handCopilot surfaces options
    Filling the bookingTypes every fieldPre-filled; agent confirms
    Holding contextIn the agent's headSurfaced, not remembered
    AI unsure or wrongFalls back to manual scheduler
    10

    Designing the AI, decision by decision

    The copilot wasn't one decision — it was a stack of them, each with options I rejected and a reason I didn't.

    Which AI approach?

    Desk research plus direct consultation with ML engineers, comparing four approaches on accuracy and speed.

    ai approach evaluation
    AI approach evaluation table comparing rule-based, NLP keyword, intent classification, and conversational LLM approaches
    Rule-based was too rigid; NLP keywords missed context; a conversational LLM was accurate but too slow. Intent classification — ~94% accuracy, fast — won; the LLM became a phase-2 bet.
    Where does the AI live?

    Four placements compared before committing.

    placement options
    Options for where the AI lives: floating widget, modal overlay, bottom drawer, and the chosen inline panel
    A floating widget broke focus; a modal killed the conversation; a drawer sat too far from the thread. The inline panel won — always visible, contextual, non-blocking.
    When should it surface a suggestion?
    surfacing decision
    Decision comparing showing everything, waiting for the agent, and the chosen auto-detect and confirm approach
    Showing everything caused overload; waiting for the agent added friction the transcript had already resolved. Auto-detect-and-confirm won: the AI detects intent, surfaces the card, the agent approves.
    Four cognitive principles

    Clarify ambiguity

    Natural language → structured intent.

    Show reasoning

    Evidence chips before acting, not a black box.

    Synthesize

    Transcript + EHR calendar + patient preferences merge into one card, not three lookups.

    Recommend, don't act

    Suggestion → script → pre-filled form, always awaiting a human confirm.

    suggestion to appointment
    Three inputs collapse into one actionable card: live call, EHR calendar, preferred days
    Three inputs — live call, EHR calendar, preferred days — collapse into one card. The EHR sync fires on its own once the agent confirms.
    ambiguity states
    Patient rejects the first slot, NextIQ re-queries availability from the transcript alone
    When the patient rejects the first slot, NextIQ re-queries from the transcript alone and dims the old suggestion with a "Mentioned" badge.
    11

    The copilot lifecycle

    I moved from low to high fidelity deliberately — flows and component logic before pixels, in a tight loop with Engineering on what the AI could reliably do.

    That process produced a live copilot lifecycle, shown across four real screens from the same call.

    01 · identify
    Call arrives already matched, transferred from the virtual assistant with caller identity in motion
    01 · Identify — the call arrives already matched, transferred with caller identity in motion before the agent speaks.
    02 · verify + understand
    Inline HIPAA verification and NextIQ intent detection at 96% confidence with evidence chips
    02 · Verify + Understand — inline HIPAA verification; then intent detection at 96% confidence with the evidence chips behind it.
    03 · act
    NextIQ drafts a suggested script and pre-fills the scheduling widget for the agent to confirm
    03 · Act — a suggested script in the agent's voice, and a pre-filled slot. Nothing sends until the agent marks it said and confirms.
    04 · resolve
    Appointment confirmed with four automations completing: form auto-filled, SMS sent, reminder scheduled, clinical note drafted
    04 · Resolve — one confirm, and the busywork completes itself: booked, texted, reminded, clinical note drafted, disposition filled.
    See it in action
    nextiq · ai copilot demo
    End-to-end walkthrough — identify, understand, act, resolve.

    Prescriptions and billing, in the same canvas. Scheduling was the deepest flow, but the intent model was built to generalize. When the patient pivots mid-call to "what do I owe?", NextIQ detects the second intent and surfaces the billing action — no new tab. Prescriptions followed the same pattern; deepening the refill flow with a live medication list was the top V2 item.

    12

    Designing for the wrong answer

    A copilot you can only trust when it's right isn't trustworthy — it's lucky.

    In healthcare, the failure path is the design, so I designed the copilot's bad days as carefully as its good ones — and every failure path lands on the manual scheduler.

    Low confidence → steps back

    Below threshold, NextIQ stops pre-filling and switches to "I'm not sure — here's what I heard," handing the agent the manual scheduler instead of a guess.

    Wrong intent → cheap to correct

    Evidence chips make a mismatch visible at a glance, before confirm; one tap opens the full editable scheduler. Correction is the fast path, not the penalty box.

    Wrong patient → the hard stop

    Intent is never acted on until inline verification matches. A mismatch escalates to the full modal and freezes the pre-fill.

    Graceful failure → no dead ends

    When the AI genuinely can't help, it says so and hands control back cleanly — to the manual path that always works.

    four-tier confidence policy
    Four-tier confidence policy: high pre-fills, medium suggests, low offers options, anything under 44% routes to a human
    The single bar agents see is the surface of a four-tier policy: high pre-fills, medium suggests with chips, low offers options with no pre-fill, and anything under 44% — or anything clinical — stays quiet and routes to a human.
    confidence indicator variants
    Four confidence indicator variants tested: bar, dots, chip, radial
    Four ways to make confidence legible — the bar shipped; agents read 96% instantly, no explanation needed.
    13

    Trade-offs I made

    Naming what I gave up — and why the trade was worth it — is how I keep a design honest.

    Verification speed vs. rigor

    Chose an inline, glanceable HIPAA check over a blocking modal on every call. Mitigated by one field at a time and a full modal for new patients and mismatches. Validated with compliance before pilot.

    Automation vs. trust

    Chose pre-fill-and-confirm over letting the AI book and send on its own. Gave up the flashier demo; one extra tap bought a system agents actually rely on (Section 14).

    Simplicity vs. flexibility

    Kept the full, editable manual scheduler behind the AI's pre-fill instead of one-click booking — because "actually, can we change that?" is the common real moment.

    14

    The disagreement I had to win

    The most important design decision on this project wasn't a screen. It was a "no."

    Midway through, a stakeholder made a reasonable-sounding push: let NextIQ book and send confirmations autonomously when confidence is high. The demo would dazzle — patient asks, appointment booked and texted, no agent tap. On paper: faster, cheaper, better in a sales deck.

    I disagreed, and had to make the case without sounding like the designer who says no to everything:

    The stakes aren't symmetric. In consumer SaaS an autonomous mistake is an annoyance; in healthcare, booking the wrong patient into the wrong specialty is a clinical and legal event. The downside was uncapped; the upside was one saved tap.

    Trust is the actual product. Agents didn't fear a slow AI — they feared being blamed for something they couldn't intercept. An agent who can't override won't lean on the tool; they'll fight it. Adoption dies quietly.

    We could get ~90% of the speed without the risk. Pre-fill-and-confirm removes the searching and typing — the real time sink — while keeping the human on the trigger.

    How I made it land: I put the autonomous version in front of agents and watched them hesitate to trust a system that could act without them — then watched the same agents relax and rely on pre-fill-and-confirm the moment they saw how easily they could override it. Confidence in the action taken rose from 3.1 to 4.6 out of 5 once override was obviously easy. Control didn't slow agents down; it was the thing that let them speed up.

    Resolution: we shipped pre-fill-and-confirm, and copilot, never autopilot became a stated pilot principle — not because I won an argument, but because the evidence did.

    15

    What testing proved

    The pivot wasn't a hunch. I re-ran the same moderated protocol on the copilot (n = 10 agents, matched tasks) against the manual baseline.

    MeasureV1 · ManualFinal · CopilotΔ
    SUS usability score6186+25 (OK → top decile)
    Unaided task completion4 / 1010 / 10+6
    Median time-on-task (book a follow-up)4m 20s1m 30s~65% faster
    Holds initiated during task7 / 100−7
    Apps / screens touched~51 canvas−4
    Cognitive load7439−35
    Agent confidence in action (1–5)3.14.6+1.5

    Illustrative quotes below — representative of session feedback, not verbatim transcripts.

    Trust & control

    "The moment I saw I could edit it before it went out, I stopped double-checking everything by hand."

    ★★★★★

    Agent 05

    Front-desk, Practice A

    Listening, not typing

    "I actually got to just talk to the patient. It was pulling the slots while I was still listening."

    ★★★★★

    Agent 09

    Front-desk, Practice C

    Speed

    "No more putting them on hold to go dig around. It's right there, already sorted."

    ★★★★★

    Agent 02

    Front-desk, Practice B

    16

    Impact — measured and projected

    It shipped as an MVP pilot across 3 practices and ~40 agents. Twelve-month production metrics don't exist yet, and I won't manufacture them.

    Measured in the study

    SUS 61 → 86; unaided completion 4/10 → 10/10; time-on-task down ~65%; cognitive load 74 → 39; and stakeholder validation — Product and clinical backed the reframe, and compliance signed off on the condensed verification, clearing it for pilot.

    Projected for the pilot

    Targets, not production results — modeled conservatively on published benchmarks, placeholders to replace with real pilot data.

    MetricBaselineTargetWhy it matters
    Average Handle Time (AHT)↓ 20–30%The 65% in-session gain discounts heavily in production
    First-Call Resolution (FCR)+15–18 ptsFewer holds + one canvas = more one-call closes
    Callback / repeat-contact rate↓ (track)Directly attacks the ~60% unanswered-call problem
    Copilot acceptance vs. override rateTrackThe trust thesis, quantified
    HIPAA verification completion100%Guardrail on the verification trade-off
    17

    Reflection

    My job wasn't to design the smartest AI in the room — it was to design the relationship between a stressed human and a fallible machine, where being wrong isn't an option. I made the case for a pivot when a tidy first version tested well enough to ship and not well enough to matter, held the line on human control against a push for more automation, and — because it's healthcare — kept the manual scheduler alive underneath so the AI never became a single point of failure.

    A V2 starts with the real practitioner calendar agents asked for, a deepened refill flow with a live medication list, and billing made permanently visible.

    The principle carries into everything I design now: for enterprise AI in healthcare, copilot — never autopilot.

    Zoomed screenshot