Isolated assistants, not a workforce
SullyAI’s vision required receptionists, nurses, scribes, coders, and clinical assistants collaborating as one team, not six disconnected tools
MaxLevels, an AI and product engineering company, partnered with Sully AI, a Y Combinator-backed healthcare AI company, to design and build a coordinated team of autonomous AI agents — receptionist, triage nurse, scribe, clinical consultant, medical coder, and care coordinator — that runs inside a health system’s existing EHR rather than beside it. The result: 12.5M+ minutes of clinical documentation automated, 59M+ minutes added back to workforce capacity, a 21x return on agent spend, and 2x provider efficiency.
6 AI Agent
Built six purpose-specific AI agents plus an orchestration layer, not one all-purpose assistant, so each agent could be tuned, audited, and scaled independently
HIPAA
HIPAA-ready architecture with human-in-the-loop review on every clinical output before it reaches the EHR
50+ EHR
Native integration with Epic and athenahealth; the platform now connects to 50+ EHR and practice-management systems
12.5M+
12.5M+ minutes of documentation automated, 59M+ minutes of capacity returned to health systems, 21x ROAS, 2x provider efficiency
AGENT SPEND
21x
Return on agent spend
MINUTES
12.5M+
Minutes of documentation automated
WORKFORCE
59M+
Minutes added back to workforce capacity
PROVIDER
2x
Provider efficiency
SullyAI’s vision required receptionists, nurses, scribes, coders, and clinical assistants collaborating as one team, not six disconnected tools
Healthcare conversations are unpredictable and specialty-specific; documentation and coding had to stay accurate across different physician workflows without slowing anyone down
Reliable reasoning across thousands of daily interactions required prioritization logic and evidence collection that stayed auditable at volume
Handling PHI meant every AI interaction needed auditability, role-based access, and human oversight, not just accuracy
Health systems couldn’t tolerate a rollout that forced providers to change how they already deliver care
The Orchestration Layer shares patient context across every agent and routes handoffs, so nothing gets re-entered and nothing falls through between the front desk and follow-up.
Receptionist handles scheduling, appointment confirmations, and intake before the patient ever reaches the front desk.
A few technical choices shaped why this platform holds up at enterprise scale.
Two model providers, routed by task, not one model for everything. Not every agent needs the same horsepower. A scheduling exchange and a real-time clinical note carry very different accuracy and latency requirements. Running OpenAI and Llama models side by side, orchestrated through LangChain, let each agent use the model suited to its job instead of forcing one model to be adequate at all six.
Docker-isolated agents. Each agent runs as its own containerized service. That’s what makes it possible to retune the coder without touching the scribe, or roll back one agent’s update without redeploying the whole platform, the same principle that keeps a large engineering team from becoming a bottleneck as the product grows.
Grafana across the pipeline. Once AI is drafting clinical documentation, visibility into what each agent is doing, and where it’s uncertain, matters as much as the automation itself. Full observability turns the audit trail into something reviewable in real time, not a log file checked after a problem surfaces.
PyTorch and TensorFlow for continuous evaluation. Clinical accuracy isn’t a one-time benchmark; it has to hold up as models change and specialties vary. SullyAI’s own published benchmarks show its agents reaching 61.17% accuracy on healthcare-specific evaluations against GPT-5’s 53.55%, exactly the kind of gap that comes from purpose-built evaluation pipelines, not general-purpose models applied to medicine as an afterthought.
SullyAI’s founders built on the premise that a hospital visit isn’t one conversation with one risk profile, it’s a sequence of distinct roles, each with different accuracy requirements and different consequences for getting it wrong. Splitting the work let each agent be tuned for its exact task instead of asking one general model to be adequate at all of them. It also kept oversight legible: a receptionist agent and a coding agent produce very different audit trails, and reviewing them separately is far simpler than untangling one system that does everything at once. And because each agent is its own module, the platform can add a new one, SullyAI has since expanded into an AI pharmacist and an AI interpreter on the same architecture, without redesigning what came before it.
Every agent operates inside a HIPAA-ready architecture with secure PHI handling, role-based access, and full auditability. A drafted note, a suggested code, a triage recommendation, none of it enters the patient record without a licensed provider reviewing it first. That review requirement wasn’t a constraint bolted on afterward; it shaped how the AI itself reasons, producing output that reads as a second opinion a clinician can check quickly, not an instruction to follow blindly.
Implementing ISO 27001 information security controls across EHR systems with real-time monitoring and compliance reporting.
Delivering continuous monitoring and automated evidence collection for SOC 2 Type II audits across different EHR environments.
Ensuring HIPAA compliance through encryption, access controls, and audit logging across EHR integrations.
Maintaining comprehensive data governance controls for Personal Data Law compliance across EHR systems.
Providing GDPR-compliant data processing with automated data mapping, consent management, and data subject request tools for EHR systems.
Meeting Canadian privacy requirements through automated privacy assessments and consent tracking across EHR platforms.
“We’re getting paid faster, our compliance risk is much lower, and our administrative overhead is down. Most importantly, we have a happier, more effective clinical team that’s providing better care to more engaged patients. Sully has become a major competitive advantage for us in the marketplace.”
21x
Return on agent spend
The platform pays for itself many times over in reclaimed clinician capacity
12.5M+
Minutes of documentation automated
Millions of hours of after-visit charting removed from clinicians’ days
59M+
Minutes added to workforce capacity
Health systems gained meaningful staff-equivalent time without adding headcount
2x
Provider efficiency
Providers see more patients in the same working day, with less administrative drag
| Matrix | AI Scribe | Sully AI |
|---|---|---|
| Scope |
|
|
| Governance |
|
|
| Integration |
|
|
01
A broader agent roster. SullyAI has extended the original six-agent team with an AI pharmacist, AI interpreter, and AI researcher, all on the same orchestration architecture.
02
Deeper EHR reach. The platform has expanded past Epic and athenahealth into Cerner (Oracle Health), MEDITECH, Allscripts, Veradigm, and 50+ systems total, plus a direct embedding inside Epic Haiku.
03
A published trust framework. SullyAI’s “Consensus Mechanism” whitepaper formalizes how its agents reach verifiable agreement with each other, a protocol for multi-agent decision-making that extends the same auditability principle this platform was built on.
04
Continuous benchmarking. Ongoing evaluation against frontier models keeps the specialty-specific accuracy gap the platform was built to create from eroding as general-purpose AI improves.
A multi-agent AI healthcare platform uses several purpose built AI agents, each handling one clinical or administrative role and coordinated by an orchestration layer that shares patient context between them. Unlike a single general-purpose assistant, each agent can be tuned, audited, and scaled independently for its specific task.
MaxLevels built a six-agent autonomous AI workforce and orchestration platform operating inside healthcare systems and EHRs to automate intake, triage, clinical documentation, decision support, and coding.
Through an orchestration layer that maintains a single shared patient context and routes handoffs between agents. When the triage nurse agent collects symptoms, the scribe and clinical consultant agents already have that context, nothing is re-entered, and nothing is dropped between the front desk and follow-up.
12.5M+ minutes of clinical documentation automated, 59M+ minutes returned to health system workforce capacity, a 21x return on agent spend, and 2x provider efficiency. The platform is now used across 100+ healthcare organizations and 30,000+ providers.
An AI medical scribe automates one step of a visit documentation. A multi-agent platform covers intake, triage, documentation, clinical decision support, coding, and follow-up as one connected workflow, with a human checkpoint on anything that reaches the patient record.
MaxLevels builds HIPAA-ready architectures with encrypted PHI handling, role-based access control, and full audit logging on every agent interaction. Human-in-the-loop review is architectural, not optional: no AI-generated note, code, or recommendation enters the patient record without licensed provider approval.
No. Every agent output, a drafted note, a suggested code, a triage recommendation passes through human review before it becomes part of the patient record. These platforms reduce administrative load; clinical decision-making authority stays with the provider.
MaxLevels built native Epic and athenahealth integration for Sully AI. The platform now connects to 50+ EHR and practice-management systems including Cerner (Oracle Health), MEDITECH, Allscripts, and Veradigm, plus a direct embedding inside Epic Haiku.
Sully AI's platform runs on React.js, Node.js, and Python, with OpenAI and Llama models routed by task through LangChain, PyTorch and TensorFlow for continuous clinical accuracy evaluation, Docker for per-agent isolation, and Grafana for pipeline-wide observability, deployed on AWS.
Yes. MaxLevels works as a healthcare AI agent development partner covering architecture, agent design, EHR integration, model evaluation, and compliance as a dedicated team or a full build partner.