Plan: automated analysis of open text

Status: draft, 2026-09-25

Goal

Reading and coding free-text answers is the most expensive manual step in the whole tool. A study with 200 participants and five open questions produces a thousand answers that a researcher reads one at a time and sorts into themes by hand. Everything else UX Testbench does is already automated; this is not.

This plan adds that automation on a two-tier architecture: a System One model (Jev) for the per-answer decisions, and a reasoning LLM (Claude Opus 5) for the judgment that decides what the answers should be sorted into. Neither runs unless a researcher turns it on for their project.

Why two tiers, and why these two

The two jobs are different in kind, not only in size.

Classifying one answer against a fixed codebook is a structured decision with a small answer set: which theme, how confident, which words support it. That is exactly what System One models are for. Jev returns a typed value and a probability distribution rather than a string, so there is no parse step and no risk of a label that is not in the codebook. Its published input price is $42 per billion input tokens with output tokens free, against $5 per million input on Claude Opus 5.

Deciding what the themes are is the opposite: open-ended, needs to weigh a whole sample, and a researcher will read and edit the result. That is a reasoning model's job, run once per study rather than once per answer.

The tier that matters most is the one below both.

TierWhat it doesRuns onCost
System ZeroRules. Empty answers, "n/a", exact and near duplicates, answers already matching a code verbatim, and the PII scanPlain Python, no networkFree
System OnePer-answer: which theme, confidence, supporting quoteJev (typesafe-sdk)Fractions of a cent per study
System TwoBuild and revise the codebook; resolve what System One could notClaude Opus 5Under a dollar per study

On real survey data System Zero clears 20–30% of rows before anything leaves the machine. The cheapest call is the one you do not make, and this is also where the privacy gate lives.

Be honest about what the split buys

It is not mainly cost. A thousand answers costs well under a dollar either way. The reasons that actually hold:

  1. Latency. A thousand answers through a reasoning model is minutes; Jev's published latency is 70–500 ms per call, so the same work finishes while the researcher watches.
  2. Consistency. The same answer should get the same code every time. A fixed codebook and a calibrated classifier give that; re-reasoning each row does not.
  3. Review surface. The researcher checks twelve themes, not a thousand decisions.
  4. Calibrated confidence. This is the one that makes the architecture work at all, and it gets its own section.

Against that, two costs to accept: prompt caches are model-scoped, so the tiers cannot share a cached prefix, and there are now two vendors in the request path instead of one.

Confidence is the load-bearing part

Escalation from System One to System Two needs a trustworthy signal for "I am not sure". An LLM's self-reported confidence is not that signal: it is generated text, and it is generated by the same process that produced the answer.

Jev reports confidence (0 to 1) and the full probabilities distribution, computed from the distribution's shape rather than asked for in words. TypeSafe's own guidance is a three-band split — under 0.5 route to a human, 0.5 to 0.9 proceed with care, above 0.9 act automatically — with the caveat that the right thresholds depend on the domain and must be measured. We measure ours (see Phase 4) rather than adopting those numbers.

There is a fit here worth naming. This project already argues, in docs/statistics.md, that a tool should report uncertainty instead of manufacturing confidence. A model trained for calibrated probabilities belongs in it. A model that always sounds sure does not.

What escalates

  • confidence below the measured threshold for that question type.
  • The answer matches no code. The strongest signal the codebook is incomplete, and the one worth acting on: it means the taxonomy missed something real.
  • The top two probabilities are within a small margin of each other.
  • A researcher marks a row miscoded. This escalates as a codebook revision task, not a re-classification. Corrections flow up into the taxonomy instead of dying as one-off edits.

The privacy problem

The landing page promises zero tracking and complete data ownership. The privacy policy makes participants a promise about where their words go. Sending free text to any external API breaks that promise unless this is built deliberately:

  • Off by default. Opt-in per project, never per instance.
  • Enabling it amends that study's consent text, so participants are told before they type, not after they have answered.
  • Identity never leaves. Send the answer text and an internal row id. Never the participant identity, email, or study slug.
  • System Zero scans for PII first — names, emails, phone numbers, ID numbers — and holds matching answers back for the researcher to decide on.
  • A local provider is a first-class implementation, not an afterthought. A university or government client that cannot send data offshore must still be able to use this. That is also the answer to depending on an early-access product for a core feature.

Two external vendors now means two data-processing agreements and two entries in the privacy policy. That is a real cost of this plan, and it belongs in the decision record.

Architecture

testbench/ai/
  __init__.py       feature flags, per-project settings
  provider.py       the seam: SystemOne and SystemTwo protocols
  jev.py            System One via typesafe-sdk
  claude.py         System Two via anthropic
  local.py          both tiers against a local model, for offline instances
  rules.py          System Zero: dedup, PII scan, trivial answers
  codebook.py       build, revise, version, diff
  coding.py         the pipeline and the escalation rules

Two narrow protocols, so no tier is welded to a vendor:

python
class SystemOne(Protocol):
    def classify(self, rows: list[TextRow], codebook: Codebook) -> list[Coded]: ...

class SystemTwo(Protocol):
    def propose_codebook(self, sample: list[str], context: StudyContext) -> Codebook: ...
    def resolve(self, rows: list[TextRow], codebook: Codebook) -> list[Coded]: ...

Jev's questions map onto the primitives directly, which is what makes the seam cheap:

Our needJev primitive
Which theme does this answer belong toChoice over the codebook
How strongly does it express the themeScore over ordered levels
Does it mention a specific thing we are trackingNoul (returns a probability)

Choice and Score carry confidence and probabilities; Noul does not, so a Noul-only question cannot drive escalation on its own.

Storage: a new text_codes table in each project's own SQLite file, keyed to the answer it codes. The original text is never overwritten and every coded answer stays one click from what the participant actually wrote. Codebooks are versioned, so a re-run is comparable to the run before it.

Execution

Phase 0 — Confirm the ground truth (half a day, blocking)

Everything below assumes access and behaviour we have read about but not verified.

  • [ ] Get early-access credentials at console.typesafe.ai; confirm Jev is usable from Indonesia and what the rate limits are.
  • [ ] Read docs.typesafe.ai/sdk/python.md and /api.md end to end. The SDK is typesafe-sdk, reads TYPESAFE_API_KEY from the environment, and exposes TypeSafeClient / AsyncTypeSafeClient with client.system_one(state=..., questions={...}). Verify the response shape against a real call rather than against this plan.
  • [ ] Settle how a thousand separate answers are sent. The documented concurrency story is AsyncTypeSafeClient plus asyncio.gather; there is no documented multi-document batch endpoint. Confirm this and measure achievable throughput.
  • [ ] Read /model-jaggedness/jev-1.13.md. Known limitations decide what we do not route here.
  • [ ] Read /legal.md and get the data-processing terms. This gates Phase 3.

Two claims to hold at arm's length until measured: "0% hallucination" is really a guarantee of type safety — a schema-valid label can still be the wrong label, and only our own eval will say how often — and the headline speed and cost multiples come from TypeSafe's own workflow evals, not ours.

Phase 1 — The seam and System Zero (2 days)

  • [ ] testbench/ai/provider.py with both protocols and a null implementation, so the app runs unchanged with no keys set.
  • [ ] testbench/ai/rules.py: normalise whitespace, drop empties and "n/a", group exact and near duplicates, PII scan.
  • [ ] Settings: AI_ENABLED, AI_SYSTEM_ONE, AI_SYSTEM_TWO, per-project opt-in in the Studio.
  • [ ] Tests for System Zero with no network at all. It is pure functions; it should be the best-tested part of this.

Phase 2 — Study-design review (2 days)

The first thing that ships, because it touches no participant data and so needs no consent change: a "Review this study" button in the Studio that reads the module YAML and flags leading questions, double-barrelled items, missing "prefer not to say", unbalanced scales, and labels that do not match their point count.

  • [ ] System Two call over the YAML; findings shown inline in the Studio editor.
  • [ ] Noul per check via System One where the check is a yes/no over one question — cheap enough to run on every save.
  • [ ] Never block a save. Advice, not a gate.

This phase proves the seam, the settings and the UI with nothing at stake.

  • [ ] Consent plumbing: enabling analysis amends consent, and a study that has already collected answers warns that existing participants did not agree to it.
  • [ ] codebook.py: System Two proposes 8–15 themes with definitions, boundary cases and example quotes from a sample.
  • [ ] The researcher edits the codebook before anything is applied. Not optional. They own the taxonomy; the model proposes it.
  • [ ] coding.py: System Zero, then System One over the remainder, then escalation to System Two by the rules above.
  • [ ] Results in the module report: theme, count, confidence distribution, example quotes, and a visible uncoded bucket. An analysis that hides what it could not classify is worse than none.
  • [ ] Re-run against a revised codebook; diff two runs.

Phase 4 — Measure it (3 days, not optional)

Routing between tiers is faith-based without this, and confidence thresholds cannot be picked from a blog post.

  • [ ] Build the eval from a study a researcher has already hand-coded. That is a real labelled set, and it costs nothing to collect.
  • [ ] Per-theme agreement between System One and the human labels.
  • [ ] Sweep the confidence threshold: at each value, what fraction escalates and what fraction of escalations the human agreed needed escalating. Pick the knee, per question type.
  • [ ] Compare against the boring baseline — System Two alone over everything. If the cascade is not clearly better on latency or agreement, say so and drop a tier. That is a valid outcome of this phase.
  • [ ] Publish the numbers in the docs. A tool that argues for statistical honesty should hold its own features to it.

Phase 5 — Card sorts and the local provider (1 week)

  • [ ] Cluster open card-sort labels ("Money stuff", "billing", "Payments & invoices") with Choice against a proposed grouping. Reuses the whole Phase 3 pipeline.
  • [ ] local.py against an Ollama-hosted model, so an instance can run all of this with nothing leaving its network. Slower and less accurate; the eval from Phase 4 says how much.
  • [ ] ADR 006: two tiers, two vendors, what participants are told, and what a local-only deployment gives up.

Deliberately not in this plan

Drafting the findings narrative. Handing a model the results and asking what they mean is where this would quietly undo what makes the project distinctive. docs/statistics.md exists to argue against exactly the confident prose an LLM produces by default. If it is ever built, the verdict, p-value and pair count are inputs the model may not contradict, and "not enough evidence yet" must survive into the text unchanged.

Participant-facing adaptive probes. Generating a follow-up question in response to what someone just typed changes the instrument mid-study. Every participant then answers a slightly different survey, which breaks comparability and is hard to consent to honestly.

LangChain or LangGraph. This project has three dependencies. Both SDKs here are small and direct, and the provider seam already gives what a framework abstraction would. Revisit only if the codebook-revision loop grows real state — and measure the plain Python loop first.

Open questions

  • Does Jev handle Indonesian well? The example study is locale: id, and nothing published says what languages the model covers. Phase 4's eval must be run on Indonesian text before this ships to any Indonesian study.
  • Who is the controller for text sent to two vendors, and does the current privacy policy survive the addition without rewriting?
  • Does an existing study's data get analysed at all, given its participants consented before this feature existed? The conservative answer is no, and it is probably the right one.
  • What happens when a researcher disagrees with the codebook wholesale? Editing 12 themes is fine; rebuilding from scratch should be one button, not a fight with the tool.

Source: docs/plans/text-analysis.md

Back to the home page