For model builders

Environments built where care happens.

Clinical workflows from health systems outside the US and Europe, run in place, with grading you can trust.

What is inside

Everything a model needs to attempt real clinical work, and everything you need to score it.

  • TasksMulti-step clinical work drawn from real practice.
  • Tools and stateThe record system, orders and messages a clinician would work with.
  • GradersRule-based and clinician-reviewed scoring, documented for each task.
  • A private held-out setFor measurement that has not leaked into training data.
Access

Three ways to run a model against an environment.

  1. In place

    Your model runs inside the health system's boundary, next to the data.

  2. Mirrored

    A de-identified or synthetic copy of the workflow, for models that can only be reached remotely.

  3. Results only

    We run your model and return scores and failure analysis.

The access mode for each environment is agreed with the health system and depends on local law.

What you get back

Results you can act on.

  • Scores by task type
  • The failures that would have changed a clinical decision, and in which direction
  • Results by language and by channel, not as one average
  • A record of how each task was written, reviewed and graded
First environment
Maternal health · Primary care

Maternal triage

Patient intake and risk assessment during pregnancy, graded against clinical guidelines.

In development
What comes next

Tell us what you cannot test.

We are choosing the next environments with our early partners. If there is a clinical workflow or a region you need, we want to hear it.