Field note / Healthcare + applied AI

Before AI writes the note,
it needs to know what belongs.

Where Jev could help organize a conversation, choose reusable components, and recognize when more context—or a better template—is needed.

By 7 min read

The idea in one line

Place the evidence. Choose the component. Preserve the meaning.

There’s a more interesting question behind ambient documentation than “Can AI write a SOAP note?”

Can it recognize what belongs, choose the right building blocks, and notice when something is missing—without filling the gaps with plausible text?

That is where I would investigate TypeSafe’s Jev. Not as another scribe, but as a decision-making step between the conversation and the assembled note. The opportunity I want to test is a note that is easier to review, and a template library that becomes easier to improve.

Give the model a smaller job

Jev works differently from a writing assistant. Its documented interface accepts text or structured text and returns bounded decisions: Choice selects an option, Noul estimates the probability of a yes/no answer, and Score rates something against a defined rubric. It does not transcribe audio or write the finished note. 1 2

That suggests a useful division of labour. A transcription service captures the conversation. Jev helps classify the evidence and evaluate candidate components. Ordinary software handles exact values and assembly rules; a writing model can help with phrasing where needed. The clinician reviews the result.

TypeSafe publishes examples of choosing among extracted text spans, shortlisting reusable skills, and checking whether a source supports a claim. Those are relevant building blocks—not evidence of clinical performance. 3 4 5

First, understand the statement—not just the speaker

SOAP organizes a note into Subjective information, Objective findings, the clinician’s Assessment, and the Plan. AWS HealthScribe already documents SOAP outputs and links summary sentences to supporting transcript segments. Sectioned, traceable notes are therefore a baseline to compare against, not an entirely new capability. 6

For this design, I would work with small, source-linked statements while preserving the surrounding dialogue. A single speaking turn might contain a symptom, an observation and a next step. Forcing the whole turn into one category loses that distinction.

Here is a synthetic example of the decisions I would want the system to make. These are illustrative expectations, not Jev test results or treatment advice.

One synthetic exchange. Four different documentation decisions.
What was said Where it belongs What the component must preserve
Patient: “My right knee—sorry, left knee—has hurt for three days, especially on stairs.” Subjective: symptom history. The correction to left, duration and activity-related qualifier.
Clinician: “On examination, the left knee is tender. There is no swelling.” Objective: focused examination. The two stated findings—not an entire pre-filled “normal knee exam.”
Clinician: “This may be an overuse problem.” Assessment: clinical impression. May be. Do not turn a provisional impression into a confirmed diagnosis.
Clinician: “Review in two weeks. Consider imaging only if it isn’t improving.” Plan: follow-up and conditional next step. The condition. This does not authorize an imaging order or book an appointment.

Speaker labels help, but they are not enough. A clinician repeating a patient’s history is not reporting a new examination finding. A patient asking for a test is not establishing the clinician’s plan. And an unanswered question is not a negative finding.

I would allow “mixed,” “not relevant” and “unclear” outcomes, and keep each statement connected to the problem it concerns. The job is to organize evidence—not infer a diagnosis the clinician never stated.

Choose components, not one oversized template

For that encounter, I would look for a symptom-history component, a focused examination, an impression and a follow-up component. Several can be appropriate at once.

Jev’s Choice selects one option, so I would not use a single Choice to select everything a note needs. I would shortlist candidates, evaluate each component’s fit, and use a single-choice decision only where the alternatives really are mutually exclusive. “None fits” must remain a valid result. 1

The closest TypeSafe example is its skill-suggestion workflow: narrow the candidates, inspect the shortlisted descriptions more closely, and allow rejection. My proposed adaptation is to select documentation components rather than software skills. 4

Each component would need more than a name: its purpose, required inputs, allowed variations and conditions that rule it out. A heading called “normal examination” should not smuggle in findings that nobody supplied.

A reusable component should carry structure—not undocumented clinical facts.

A mismatch can mean four different things

This is where I think the idea becomes more valuable than a macro trigger. Once the system has a candidate, it should distinguish why the component does not quite fit.

Modify it. The component is appropriate, but needs a supported qualifier: left rather than right, intermittent rather than constant, suspected rather than confirmed. Jev could classify the qualifier or select its source span. The application should preserve the actual evidence, not ask the model to invent a missing value. TypeSafe’s extraction example explicitly separates finding candidates, selecting one and copying its value. 3

Add a component. A separate issue or follow-up instruction needs its own place. Add a supported building block rather than stretching the first one to cover everything. An addition should represent something stated or otherwise verified—not what a typical encounter usually includes.

Get more context. “It has been worse since then” may require the preceding exchange. “Continue the same dose” may require reconciliation with the medication record. I would request the specific missing context or ask for clarification. “Not discussed,” “not heard clearly” and “conflicting evidence” should not all become “normal.”

Flag the template for refactoring. If clinicians repeatedly strip unrelated findings out of a bundled examination, the reusable component may be too large. The answer might be smaller, independently selectable sections. Jev could help flag and categorize those mismatches; it would not write the refactor or publish a new clinical template itself.

These are not necessarily exclusive. A component may need both a modifier and clarification. One overall “fit score” would hide the useful detail.

Fix this note now. Improve the library separately.

I would keep two loops distinct.

The first handles the current encounter: place the evidence, select the components, preserve qualifiers and resolve meaningful gaps.

The second reviews patterns across approved edits: are people correcting transcription errors, rejecting the wrong component, expressing a writing preference, or compensating for a badly structured template?

Those are different product problems. A repeated edit is a clue, not automatic permission to change everyone’s workflow. A clinical owner should approve library changes, version them and check existing examples before release. Reorganizing a component without changing its meaning is a refactor; changing its clinical content is a separate decision.

There is no need to suggest that Jev automatically learns a practice’s preferences. TypeSafe says the same model weights serve its customers; domain adaptation comes through supplied context, questions and criteria. The learning loop I am proposing belongs to the product and its owners. 2

More context is not always better context

I would give Jev the relevant exchange, speaker and time information, the candidate component’s requirements, and only the chart context needed for that decision. Historical chart data should stay distinguishable from today’s conversation.

TypeSafe warns that irrelevant context, numeric reasoning, indirect questions and adversarial content can undermine Jev 1.13. That argues for focused questions and exact calculations in code—not an entire chart attached to every request. 7

The model could recommend a context category to retrieve, but the application should control access. Independent questions can run together; a question that depends on newly retrieved information needs a later evaluation. TypeSafe documents that questions in one request are evaluated independently against the same supplied state. 8

For the clinician, I would present a simple proposed change and its supporting excerpt. Highlight an unresolved issue when it matters. A wall of confidence percentages would not be a better review experience.

The test is less correction—not more automation

I would compare three approaches on the same clinician-reviewed cases: the existing template workflow, a general language model, and a hybrid using Jev for the smaller decisions. Start with one encounter type and synthetic or appropriately approved data.

The useful measures would be important facts omitted, statements placed incorrectly, unsupported findings inserted, qualifiers lost, inappropriate components selected, unnecessary clarification requests, and total clinician review time. Cost should include transcription, retrieval, every model call and rework—not just Jev’s API charge. I would also test later corrections, unclear audio, family-member histories and multiple problems in one conversation.

Confidence is useful for deciding what deserves attention. It is not a clinical accuracy certificate. TypeSafe derives it from the answer distribution; early independent Jev research also shows why thresholds and consistency across related questions need testing. Neither those general benchmarks nor the vendor examples establish this SOAP application. 9 10 11

ACI-BENCH provides public dialogue–note examples that could help seed an evaluation, but component selection, modifier preservation and refactor decisions would need additional clinician-reviewed labels. 12

I would keep suggestions out of the signed record until reviewed, and settle the patient-data arrangements before any live evaluation. If the extra decision step does not reduce correction burden or improve fidelity, I would leave it out.

My bet is not that Jev should write the note. It is that a better-organized set of decisions could make the note easier to trust—and reveal which parts of the documentation workflow deserve redesign.

Evidence behind the idea

Primary documentation and original research reviewed on 8 October 2026. The proposed workflow, synthetic example and evaluation criteria are Jay Sethi’s synthesis. No Jev API calls, clinical pilot or patient-data processing were performed for this note. The sources below establish relevant capabilities and limitations, not the clinical effectiveness of this application. Cookbook results use their own model versions and nonclinical tasks; their performance is not transferred to SOAP notes here.

  1. TypeSafe · API reference

    Choice, Noul and Score; Choice returns one option. These are documented interfaces, not clinical validation.

  2. TypeSafe · Models

    Current documentation identifies Jev 1.13.0 as text-only and explains request-based domain configuration rather than per-customer fine-tuning.

  3. TypeSafe · Pre-parsed value extraction

    A vendor example separating candidate discovery, selection and exact copying. Its worked examples are not clinical encounters.

  4. TypeSafe · Skill suggestion

    A vendor example of shortlisting and re-checking reusable skills, including rejecting all candidates. The proposed clinical adaptation is the author’s.

  5. TypeSafe · Double-checking citations

    A vendor example of checking source presence and whether the surrounding context supports a claim. It does not establish clinical correctness.

  6. AWS · HealthScribe clinical documentation

    Documents physical- and behavioural-health SOAP templates and sentence-level links to transcript segments. A market baseline, not a Jev capability.

  7. TypeSafe · Jev 1.13 limitations

    Version-specific cautions about irrelevant context, number/date reasoning, indirect questions, option order and adversarial content; reviewed by the vendor 2 October 2026.

  8. TypeSafe · State

    Each question evaluates the same supplied state independently. New context must be supplied explicitly; dependent steps are not an automatic reasoning chain.

  9. TypeSafe · Confidence

    Confidence is computed from answer probabilities. Noul has no separate confidence field. Thresholds require use-case evaluation.

  10. Deußer, Sparrenberg & Sifa · Evaluating and Benchmarking the System One Model Jev

    September 2026 preprint reporting general-task evaluations and threshold sensitivity. Not a SOAP-routing or clinical deployment study.

  11. Li, He & Li · Beyond Calibration

    September 2026 preprint examining consistency across logically related questions. Includes PubMedQA items, not a validation of ambient note assembly.

  12. Yim and colleagues · ACI-BENCH

    Scientific Data, 2023. Public dialogue–note benchmark; additional annotation would be needed to evaluate the proposed component and modification decisions.

Make the next product decision clearer.

Jay Sethi · Product strategy, experience architecture & applied AI. Research, workflow definition and practical direction for complex products.

Discuss a role or project ↗