The Sully AI platform shown on a tablet, with patient data populating into an EHR

How MaxLevels Built Sully AI's Multi-Agent Healthcare Workforce for Health Systems

MaxLevels, an AI and product engineering company, partnered with Sully AI, a Y Combinator-backed healthcare AI company, to design and build a coordinated team of autonomous AI agents — receptionist, triage nurse, scribe, clinical consultant, medical coder, and care coordinator — that runs inside a health system’s existing EHR rather than beside it. The result: 12.5M+ minutes of clinical documentation automated, 59M+ minutes added back to workforce capacity, a 21x return on agent spend, and 2x provider efficiency.

  • 6 AI Agent

    Built six purpose-specific AI agents plus an orchestration layer, not one all-purpose assistant, so each agent could be tuned, audited, and scaled independently

  • HIPAA

    HIPAA-ready architecture with human-in-the-loop review on every clinical output before it reaches the EHR

  • 50+ EHR

    Native integration with Epic and athenahealth; the platform now connects to 50+ EHR and practice-management systems

  • 12.5M+

    12.5M+ minutes of documentation automated, 59M+ minutes of capacity returned to health systems, 21x ROAS, 2x provider efficiency

At a Glance

Founders
Client
SullyAI
Industry
Healthcare, Clinical and Administrative AI
Headquarters
Mountain View, California
Engagement
Dedicated hire team, AI platform build partner
Compliance
HIPAA-ready architecture with full auditability, ISO 27001 and SOC 2 Type II certifications
Services Provided
AI agent development, healthcare software engineering, EHR integration, AI/ML evaluation and data science
Backing
Y Combinator, 500 Startups, Plug and Play, Sequoia Scout Fund
Stack
React.js, Node.js, Python; OpenAI, Llama, LangChain, PyTorch, TensorFlow; AWS, Docker, Grafana

Autonomous impact
on Hospitals

  • AGENT SPEND

    21x

    Return on agent spend

  • MINUTES

    12.5M+

    Minutes of documentation automated

  • WORKFORCE

    59M+

    Minutes added back to workforce capacity

  • PROVIDER

    2x

    Provider efficiency

Why This Had
to Be Built

  • 41.9% of physicians reported at least one symptom of burnout in 2025.
  • 86,000 physician shortfall projected by 2036.
  • 34% growth expected in the 65-and-older population by that same year.
  • Sources: American Medical Association, 2025 Organizational Biopsy · AAMC, “The Complexities of Physician Supply and Demand”.
  • Burnout is actually improving, physician burnout has fallen for four straight years. But the underlying math hasn’t more patients are coming, fewer doctors are available to see them, and most of what still eats a clinician’s day, charting, coding, chasing a callback, isn’t the practice of medicine at all. A single AI scribe helps with one piece of that. It doesn’t touch the other four.

The Challenge

Isolated assistants, not a workforce

SullyAI’s vision required receptionists, nurses, scribes, coders, and clinical assistants collaborating as one team, not six disconnected tools

Clinical accuracy across real-world complexity

Healthcare conversations are unpredictable and specialty-specific; documentation and coding had to stay accurate across different physician workflows without slowing anyone down

Scale without losing transparency

Reliable reasoning across thousands of daily interactions required prioritization logic and evidence collection that stayed auditable at volume

Autonomy versus governance

Handling PHI meant every AI interaction needed auditability, role-based access, and human oversight, not just accuracy

Integration without disruption

Health systems couldn’t tolerate a rollout that forced providers to change how they already deliver care

How Sully’s Agents Work a Patient Visit

The Orchestration Layer shares patient context across every agent and routes handoffs, so nothing gets re-entered and nothing falls through between the front desk and follow-up.

STEP 1

AI Receptionist

Always on call for
scheduling, intake, & reminders

Receptionist handles scheduling, appointment confirmations, and intake before the patient ever reaches the front desk.

AI Receptionist workflow interface showing appointment scheduling and patient intake

Engineering Decisions Worth Knowing About

A few technical choices shaped why this platform holds up at enterprise scale.

Two model providers

Two model providers, routed by task, not one model for everything. Not every agent needs the same horsepower. A scheduling exchange and a real-time clinical note carry very different accuracy and latency requirements. Running OpenAI and Llama models side by side, orchestrated through LangChain, let each agent use the model suited to its job instead of forcing one model to be adequate at all six.

Docker-isolated agents.

Docker-isolated agents. Each agent runs as its own containerized service. That’s what makes it possible to retune the coder without touching the scribe, or roll back one agent’s update without redeploying the whole platform, the same principle that keeps a large engineering team from becoming a bottleneck as the product grows.

Grafana across the pipeline

Grafana across the pipeline. Once AI is drafting clinical documentation, visibility into what each agent is doing, and where it’s uncertain, matters as much as the automation itself. Full observability turns the audit trail into something reviewable in real time, not a log file checked after a problem surfaces.

PyTorch and TensorFlow for evaluation.

PyTorch and TensorFlow for continuous evaluation. Clinical accuracy isn’t a one-time benchmark; it has to hold up as models change and specialties vary. SullyAI’s own published benchmarks show its agents reaching 61.17% accuracy on healthcare-specific evaluations against GPT-5’s 53.55%, exactly the kind of gap that comes from purpose-built evaluation pipelines, not general-purpose models applied to medicine as an afterthought.

Why Six Agents,
Not One Assistant

SullyAI’s founders built on the premise that a hospital visit isn’t one conversation with one risk profile, it’s a sequence of distinct roles, each with different accuracy requirements and different consequences for getting it wrong. Splitting the work let each agent be tuned for its exact task instead of asking one general model to be adequate at all of them. It also kept oversight legible: a receptionist agent and a coding agent produce very different audit trails, and reviewing them separately is far simpler than untangling one system that does everything at once. And because each agent is its own module, the platform can add a new one, SullyAI has since expanded into an AI pharmacist and an AI interpreter on the same architecture, without redesigning what came before it.

Built for Clinical Trust, Not Just Clinical Speed

Every agent operates inside a HIPAA-ready architecture with secure PHI handling, role-based access, and full auditability. A drafted note, a suggested code, a triage recommendation, none of it enters the patient record without a licensed provider reviewing it first. That review requirement wasn’t a constraint bolted on afterward; it shaped how the AI itself reasons, producing output that reads as a second opinion a clinician can check quickly, not an instruction to follow blindly.

  • ISO 27001

    Implementing ISO 27001 information security controls across EHR systems with real-time monitoring and compliance reporting.

  • SOC 2 Type II

    Delivering continuous monitoring and automated evidence collection for SOC 2 Type II audits across different EHR environments.

  • HIPAA

    Ensuring HIPAA compliance through encryption, access controls, and audit logging across EHR integrations.

  • PDL

    Maintaining comprehensive data governance controls for Personal Data Law compliance across EHR systems.

  • GDPR

    Providing GDPR-compliant data processing with automated data mapping, consent management, and data subject request tools for EHR systems.

  • PIPEDA

    Meeting Canadian privacy requirements through automated privacy assessments and consent tracking across EHR platforms.

The Results

“We’re getting paid faster, our compliance risk is much lower, and our administrative overhead is down. Most importantly, we have a happier, more effective clinical team that’s providing better care to more engaged patients. Sully has become a major competitive advantage for us in the marketplace.”
Derek Ayers, DO Chief Medical Officer
  • 21x
  • 12.5M+
  • 59M+
  • 2x

  • 21x

    Return on agent spend

    The platform pays for itself many times over in reclaimed clinician capacity

  • 12.5M+

    Minutes of documentation automated

    Millions of hours of after-visit charting removed from clinicians’ days

  • 59M+

    Minutes added to workforce capacity

    Health systems gained meaningful staff-equivalent time without adding headcount

  • 2x

    Provider efficiency

    Providers see more patients in the same working day, with less administrative drag

AI scribe alone vs.
SullyAI’s agent workforce

Matrix AI Scribe Sully AI
Scope Not supported Documentation only Supported by Sully AI Intake, triage, documentation, coding, and follow-up
Governance Not supported One system doing every job Supported by Sully AI Each agent has a defined role and its own audit trail
Integration Not supported Often a bolt-on tool Supported by Sully AI Native inside Epic and athenahealth, plus 50+ other EHR and practice systems

What’s Next for SullyAI’s Platform

  1. 01

    A broader agent roster. SullyAI has extended the original six-agent team with an AI pharmacist, AI interpreter, and AI researcher, all on the same orchestration architecture.

  2. 02

    Deeper EHR reach. The platform has expanded past Epic and athenahealth into Cerner (Oracle Health), MEDITECH, Allscripts, Veradigm, and 50+ systems total, plus a direct embedding inside Epic Haiku.

  3. 03

    A published trust framework. SullyAI’s “Consensus Mechanism” whitepaper formalizes how its agents reach verifiable agreement with each other, a protocol for multi-agent decision-making that extends the same auditability principle this platform was built on.

  4. 04

    Continuous benchmarking. Ongoing evaluation against frontier models keeps the specialty-specific accuracy gap the platform was built to create from eroding as general-purpose AI improves.

Frequently Asked Questions

A multi-agent AI healthcare platform uses several purpose built AI agents, each handling one clinical or administrative role and coordinated by an orchestration layer that shares patient context between them. Unlike a single general-purpose assistant, each agent can be tuned, audited, and scaled independently for its specific task.

MaxLevels built a six-agent autonomous AI workforce and orchestration platform operating inside healthcare systems and EHRs to automate intake, triage, clinical documentation, decision support, and coding.

Through an orchestration layer that maintains a single shared patient context and routes handoffs between agents. When the triage nurse agent collects symptoms, the scribe and clinical consultant agents already have that context, nothing is re-entered, and nothing is dropped between the front desk and follow-up.

12.5M+ minutes of clinical documentation automated, 59M+ minutes returned to health system workforce capacity, a 21x return on agent spend, and 2x provider efficiency. The platform is now used across 100+ healthcare organizations and 30,000+ providers.

An AI medical scribe automates one step of a visit documentation. A multi-agent platform covers intake, triage, documentation, clinical decision support, coding, and follow-up as one connected workflow, with a human checkpoint on anything that reaches the patient record.

MaxLevels builds HIPAA-ready architectures with encrypted PHI handling, role-based access control, and full audit logging on every agent interaction. Human-in-the-loop review is architectural, not optional: no AI-generated note, code, or recommendation enters the patient record without licensed provider approval.

No. Every agent output, a drafted note, a suggested code, a triage recommendation passes through human review before it becomes part of the patient record. These platforms reduce administrative load; clinical decision-making authority stays with the provider.

MaxLevels built native Epic and athenahealth integration for Sully AI. The platform now connects to 50+ EHR and practice-management systems including Cerner (Oracle Health), MEDITECH, Allscripts, and Veradigm, plus a direct embedding inside Epic Haiku.

Sully AI's platform runs on React.js, Node.js, and Python, with OpenAI and Llama models routed by task through LangChain, PyTorch and TensorFlow for continuous clinical accuracy evaluation, Docker for per-agent isolation, and Grafana for pipeline-wide observability, deployed on AWS.

Yes. MaxLevels works as a healthcare AI agent development partner covering architecture, agent design, EHR integration, model evaluation, and compliance as a dedicated team or a full build partner.