Open benchmark

Measure the words, the speakers, and the failures.

Daisy will not publish a magic accuracy percentage. The benchmark is an executable protocol that keeps raw output, versions, references, and scoring rules together so every row can be checked.

Current status

Neutral scorerReadyWER, CER, DER, JER, speaker count, capture completeness and RTF.
Daisy product runnerReadyRuns the same archive decoder, final Whisper profile and FluidAudio diarizer as the app.
Synthetic pipeline smokePassedValidates the harness only. TTS numbers are never product accuracy evidence.
Public AMI baselinePublishedOne real 17-minute, four-speaker case with raw output, reference, hashes and environment.
Daisy vs Humla vs OpenWhisprMeasuringPublic results stay hidden until the shared real/public dataset is complete.

First public baseline

AMI ES2004a Mix-Headset · 17:29 · English · 4 speakers · automatic speaker count. This is one reproducible diarization case, not a claim about every meeting.

SystemDER ↓JER ↓SpeakersCaptureRTF ↓
Daisy 1.0.7.5915.68%20.28%4 / 4100%0.122×
HumlaNot measuredNot measuredNot measuredNot measuredNot measured
OpenWhisprNot measuredNot measuredNot measuredNot measuredNot measured

Words-only RTTM reference, 0.25 s collar, overlap scored. RTF is the median of three warm runs (0.139× / 0.122× / 0.107×). WER/CER stay blank until the transcript reference is normalized. Competitors stay blank until their uncorrected output exists for this exact WAV.

Open raw evidence and report ↗

What gets measured

WER / CER

Word and character errors after one shared multilingual normalization policy.

DER / JER

Speaker confusion, missed speech and false alarms. Labels are permutation-invariant; overlap is scored.

Capture

Whether microphone and system audio survived for the complete expected duration.

Time

Processing RTF plus warm p50/p95 where interactive latency matters.

Recovery

Sleep, lid close, route changes, force-quit and whether a usable archive remains.

Ownership / MCP

What exports without a vendor service and whether current clients receive usable tools.

The case queue

Every case in the matrix, with its status. Russian, English and mixed RU↔EN cases are defined and run through the same product runner; a row fills in only when its full evidence passes the publication gate below.

CaseLanguageLengthSpeakersCaptureStatus
AMI ES2004aEnglish17 min4single trackPublished
Clean conversationEnglish15 min2mic + systemNot measured
Clean conversationRussian15 min2mic + systemNot measured
Code-switchingRussian ↔ English15 min2mic + systemNot measured
Meeting with overlapEnglish60 min4mic + systemNot measured
Large callRussian60 min6systemNot measured
Long sessionEnglish180 min2mic + systemNot measured
Force-quit recoveryany40 min1micNot measured

Private cases will use recordings made with every participant's consent: the audio stays private, the scores and raw hypotheses are published.

The shared test matrix

Duration
15 / 60 / 180 minutes
Language
English / Russian / mixed RU↔EN
Speakers
2 / 4 / 6, with automatic and hinted count recorded separately
Audio
Clean / noise / overlap and cross-talk
Capture
Microphone + system audio, including failure and recovery paths
Runs
Three warm measured runs; exact Mac, OS, app, model and settings pinned

Publication gate

A result appears here only with a reference transcript, RTTM where relevant, raw uncorrected hypothesis, SHA-256, exact build and model settings, environment, and scorer output. Weak cases and unavailable features remain in the table.

Inspect the harness on GitHub ↗Protocol last updated 28 September 2026.