actAVA Logo

Pioneering Enterprise AI for Healthcare, Built by Industry Veterans.

Ready to transform your healthcare AI?

actAVA Logo

Pioneering Enterprise AI for Healthcare, Built by Industry Veterans.

Ready to transform your healthcare AI?

Contact

General

info@actava.ai

Sales

sales@actava.ai

Support

support@actava.ai

Locations

Headquarters

4695 Chabot Drive Suite 200, Pleasanton, CA 94588

Products

  • KORA
  • Agent Building + Orchestration
  • Agent Testing + Remediation
  • Agent Learning + Governance
  • CURA
  • CHRYSO
  • Available on AWS Marketplace

AI Transformation

  • Overview
  • Build & Accelerate
  • Measure & Control
  • Govern & Orchestrate

About

  • Our Technology
  • Our Team
  • Our Advisors
  • Our Investors
  • Our Partners
  • Our Customers

Compliance

  • Compliance

Solutions Library

  • Agent Workflow Library

Models

  • Cura 1T
  • API Docs

Benchmarks

  • Leaderboard
  • CHI-Bench

News

  • Blog
  • Release Notes
  • Press Release
  • Resources

Company

  • Home
  • Careers
  • Trust

© 2026 actAVA, Inc. All rights reserved.

Privacy PolicyTerms of Service

Compliance & Certification

actAVA is HIPAA compliant and certified by Delve.

HIPAA compliance badge — Monitored by DelveSOC 2 Type 1 compliance badge — Monitored by DelveSOC 2 Type 2 compliance badge
Chat with AVA
Home
Products
KORAAgent Building + OrchestrationAgent Testing + RemediationAgent Learning + Governance
CURACHRYSOχ-BENCHAvailable on AWS Marketplace
AI Transformation
OverviewBuild & AccelerateMeasure & ControlGovern & Orchestrate
Compliance
Models
Cura 1TAPI Docs
Benchmarks
LeaderboardCHI-Bench
Solutions Library
About
Our TechnologyOur TeamOur AdvisorsOur InvestorsOur PartnersOur CustomersCareers
News
BlogRelease NotesPress ReleaseNewsroomResources
Home
Compliance
Solutions Library

Healthcare IT News

Featured

Why Johns Hopkins is benchmarking AI agents before deployment

Health system leaders say reliable benchmarking, governance and workflow design must come before scaling agentic AI across administrative operations. Johns Hopkins Medicine is taking a deliberately cautious approach to agentic AI, focusing first on proving reliability and governance before expecting measurable financial returns.

July 24, 2026·Read on Healthcare IT News

Newsweek

Have health care’s AI ambitions hit a reliability wall?

As health care increasingly adopts AI, experts say reliability and governance may matter more than model size.

July 23, 2026·Read on Newsweek

AI Weekly

actAVA debuts Cura 1T, a self-evolving healthcare LLM

actAVA AI released Cura 1T, a healthcare-specialized LLM trained through what the authors call a human-gated self-evolution loop. In each round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from failures.

July 17, 2026·Read on AI Weekly

GitHub

actava-ai/Cura

Cura 1T is actAVA's one-trillion-parameter healthcare model. Fine-tuned from Kimi-K2.6 through recursive self-improvement — each round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures, with humans gating every keep-or-revert decision — it handles the work healthcare models are actually asked to do: Patient consultation: open-ended health conversations, graded hard by physician-written rubrics Clinical reasoning over text and medical images: expert-exam-level diagnosis, treatment, and basic-science questions (text + vision, 256K context) Interactive diagnosis and EHR tool use: multi-turn consultations and FHIR-based record operations, end to end

July 16, 2026·Read on GitHub

Hugging Face

Cura 1T: Specialized Model for Agentic Healthcare

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.

July 16, 2026·Read on Hugging Face

arXiv

Cura 1T: Specialized Model for Agentic Healthcare

Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health record (EHR) tool use. These capabilities fail in different ways, and a narrow update for one task can degrade another. We present Cura 1T, a healthcare-specialized LLM trained through a human-gated self-evolution loop. In each evolution round, a training agent plans a target capability, trains the model, evaluates benchmark trajectories, and refines the data mixture from observed failures. This data-centered loop improves the model through targeted synthetic and curated examples rather than a single generic medical-data update. Across the healthcare evaluation suite, Cura 1T ranks at or near the top among frontier baselines, while remaining competitive on out-of-domain reasoning and agentic benchmarks.

July 16, 2026·Read on arXiv

Postcards From the Edge

EP 620 | Daily AI News | July 15, 2026: The True Playbook for Agentic Scaling

Cura 1T is a one-trillion-parameter healthcare-focused AI model designed for clinical reasoning, patient communication, and healthcare workflows, with strong benchmark performance against existing medical AI systems. Our analysts regarded it as an impressive vertical AI model for healthcare, but noted that its impact is largely limited to healthcare organizations and that rapid advances in foundation models may quickly narrow its competitive advantage.

July 15, 2026·Read on Postcards From the Edge

Qiita

Claude, GPT, and Gemini Fail 72% of the Time in Medical Settings! The Brutal Truth Exposed by CHI-Bench: Can AI Agents Can't Be Used in Hospitals?

A benchmark released yesterday (May 20, 2026) sent shockwaves through the AI industry. Thirty AI agents — including the latest Claude, GPT-5.5, and Gemini — failed 72% of the time on U.S. medical workflows.

May 21, 2026·Read on Qiita

Carroll County News

Claude, GPT, Gemini Agents Fail 72% of U.S. Healthcare Workflows, New Benchmark Finds

AI company actAVA.ai today released CHI-Bench, the world’s first long-horizon healthcare benchmark for AI agents. Across 75 workflows and 30 frontier agents from Anthropic, OpenAI, Google, x.AI, DeepSeek, and Z.ai, the best-performing agent fails roughly seven out of ten real clinical cases. Code, data, and the live leaderboard are at actava.ai/benchmarks.

May 20, 2026·Read on Carroll County News

The Daily News (Galveston)

Claude, GPT, Gemini Agents Fail 72% of U.S. Healthcare Workflows, New Benchmark Finds

Across the 30 frontier agents tested, Anthropic's Claude Code with Opus 4.6 achieved the best overall performance at 28% pass@1, followed by OpenAI's Codex with GPT-5.5 at 21%. By domain, utilization review reached 41%, care management 32%, and prior-authorization paperwork 29%. Reliability remained a major issue, with no agent clearing 20% when the same case was run three times. Under endurance testing, where agents were asked to handle 25 cases in one session, the best system completed under 4%. In a fully end-to-end setting, where one AI submitted a prior-auth request and a second acted as the UM reviewer, no task passed successfully.

May 20, 2026·Read on The Daily News (Galveston)

FinancialContent

Claude, GPT, Gemini Agents Fail 72% of U.S. Healthcare Workflows, New Benchmark Finds

AI labs position agents as ready for long workflows, but until now no public benchmark validated that claim in healthcare, where one missed policy check can mean a denied authorization, delayed treatment, or audit finding. Each trial in CHI-Bench runs an agent for 60-80 steps across four to six clinical stages, exposing 21 healthcare apps through 200+ MCP tools and a 1,279-document operations handbook. It evaluates the trajectory, every artifact, and world state using deterministic unit tests and LLM judge for evidence grounding, consent, and cross-stage consistency.

May 20, 2026·Read on FinancialContent

GitHub

actava-ai/chi-bench

Χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? χ -Bench evaluates AI agents on end-to-end U.S. healthcare workflows across three long-horizon domains: provider prior authorization, payer utilization management, and population care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed over MCP, with a 1,279-document Managed-Care Operations Handbook skills, and asks it to drive the case through tool calls and artifact authoring.

May 15, 2026·Read on GitHub

arXiv

χ-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; multi-role composition, a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce χ -Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role’s artifacts, guided by a 1,279-document managed-care operations handbook skill. Across 30 agent harness/model configurations, the best agent resolves only 28.0% of tasks, no agent clears 20% on strict pass^3, and executing all tasks in a single session slumps the performance to 3.8%. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.

May 15, 2026·Read on arXiv