60 client emails. Four models. Every answer.

Meow builds private AI for businesses that can't let their data out, and chooses each model by testing it first. This is one such test, in full: pick an email and compare what each model answered with the answer checked by hand.

Checked by hand0.6Bin a browser1.7Bin a browser8Bin a browser32Bone server GPU32B + rulessame GPU
TopicsSalaryWrong: Salary, AccountingRight: SalaryRight: SalaryRight: SalaryRight: Salary
UrgencyLowWrong: NormalRight: LowRight: LowRight: LowRight: Low
DeadlineNoneRight: noneWrong: currentRight: noneRight: noneRight: none

The right model is measured, not guessed

Added up over the 20 emails, in each language. 0.6B, 1.7B and 8B ran in a browser on an Apple M1 laptop; 32B ran on one server graphics card, the size of model a private deployment serves to a whole team. All four got exactly the same instructions.

Triage, measured: size aloneCorrect answers out of 20 emails per language, with the same instructions for every model. Times are medians: 32B on a server GPU, the others on an Apple M1 laptop.
English
0.6B1.7B8B32B
Deadline19141815
Topics11121011
Urgency9111111
All three right6353
Time per emailMedian1.3 s2.4 s6.6 s1.5 s
French
0.6B1.7B8B32B
Deadline12122019
Topics1298
Urgency591412
All three right0053
Time per emailMedian1.5 s3.0 s6.7 s1.7 s
German
0.6B1.7B8B32B
Deadline13122018
Topics3746
Urgency9141010
All three right1432
Time per emailMedian1.5 s3.1 s7.4 s1.6 s
The same 32B, with the rules written inThe instructions above name the fields and nothing more. Here they also state the rules the answers were checked against, word for word, as they were written before any model saw the emails.
English
Same instructionsRules written in
Deadline1520
Topics1110
Urgency1119
All three right39
French
Same instructionsRules written in
Deadline1920
Topics86
Urgency1219
All three right36
German
Same instructionsRules written in
Deadline1820
Topics67
Urgency1020
All three right27
The rules, as added to the instructions

Rules: urgency is high when the client says it is urgent or a penalty or loss is imminent, low when the email is information only with nothing to do, and normal for any other request. The deadline is the date the client gives for the fiduciary to act, or none. The topics are every topic the email asks about.

What the test showed

  1. A bigger model closes only part of the language gap. In French, 8B found every deadline, 20 of 20, and named the topics right 9 times, against once for 0.6B. In German it also found every deadline, but filed 19 of the 20 emails under "Other".
  2. Bigger is not automatically better. With the same instructions, 32B got all three right less often than 8B in every language: it marked as urgent emails the checked answers rate as normal, and added topics the client didn't ask about.
  3. Once a model is large enough, instructions matter more than size. With the rules written in, the same 32B judged urgency right in 19 or 20 emails of 20 and found every deadline, in all three languages. Spelled-out urgency rules had made the smallest model call almost every email urgent.
  4. Topics are the hard part. Even with the rules, 32B picked exactly the right topics in only 6 to 10 emails of 20, mostly by adding a topic the email only mentions. That is what examples from your own emails are for.

Read the sample report

With the rules written in, the same 32B model went from 3 to 9 emails of 20 with all three right in English. What remains is judgement on topics: the next step is examples from your own emails, checked with the same test.

Book a free 15‑minute call

Agents that ask before they act

Stackive In daily use

Stackive is the platform built in-house to run the company: email, invoicing and deployments go through it, each in a service of its own, with one agent layer that works across them.

Its agents read what a task needs and prepare the work. Anything that sends an email, issues an invoice or deploys code waits for a person on the team to approve it, and every step is recorded. The rough edges show up here first, before the same kind of agents are set up for you.

Every entry names who acted, an agent or a person, and nothing leaves until a person has said yes.

A morning in the log. An illustration with made-up entries.
  1. 08:57Email agentSorted 14 new emails into client foldersDone
  2. 09:02Email agentDrafted a reply to a supplier about a late deliveryWaiting for a person
  3. 09:05A personApproved the replyApproved
  4. 09:05Email agentSent the replySent
  5. 10:31Billing agentPrepared an invoice for 12 hours of workWaiting for a person
  6. 10:40A personSent it back: wrong hourly rateSent back
  7. 10:42Billing agentCorrected the rate and prepared it againWaiting for a person
  8. 11:15Delivery agentPrepared version 4.2 for releaseWaiting for a person

Accessible software

  1. Voice search that stays on the device

    Local voice assistant Prototype, 2026

    A French voice assistant for searching and playing an audiobook catalogue hands-free, running entirely in the browser: Whisper speech recognition, a language model that turns spoken requests into catalogue searches, and speech synthesis, all on the listener's device, with no cloud speech service.

    Every model and prompt change is scored against a hand-checked corpus of real voice requests, so improvements are measured rather than assumed. Once loaded, the interface reopens without a connection.

    Runs on
    The listener's device
    Models
    Whisper, a small language model, speech synthesis
    Measured on
    Real voice requests, checked by hand
    On the listener's device
    1. Spoken question
    2. Transcribed, then searched
    3. Spoken answer
  2. An audiobook library in the browser

    Web Player Web app, 2023

    The Bibliothèque Braille Romande's catalogue as a web app: more than 14,000 audiobooks to search, listen to online or download, with a member login and help pages on braille and DAISY books.

    Built for the keyboard and screen readers first: skip links to every part of the page, and the expected screen-reader behaviour written down as testable rules, checked by automated tests with Playwright and Axe. Its audiobook player and accessibility toolbar are published as open-source packages.

    Runs in
    Any web browser
    Catalogue
    More than 14,000 audiobooks
    Checked with
    Screen-reader tests and Axe audits
    The web player's home page: search, member login and the latest audiobooks
  3. Open-source building blocks

    a11y-player and a11y-button npm packages, 2025

    The web player's audiobook player and accessibility toolbar, published as React packages so any app can use them.

    a11y-player plays DAISY audiobooks: fully keyboard-navigable, working with screen readers, with section navigation, playback speed and shareable bookmarks. a11y-button adds a toolbar where readers switch to high contrast, change text size and spacing, use a reading mask or dyslexia-friendly fonts, and keep their choices.

    Licence
    MIT
    For
    React apps
    Used in
    The Web Player
    $ npm install a11y-player a11y-button
    
    import { DaisyPlayer } from 'a11y-player';
    import {
      AccessibilityProvider,
      AccessibilityToolbar,
    } from 'a11y-button';
    
    <AccessibilityProvider>
      <AccessibilityToolbar />
      <DaisyPlayer
        dirUrl={bookUrl}
        appUrl={appUrl}
      />
    </AccessibilityProvider>
    
  4. Existing apps, rebuilt for blind readers

    BBR Player™ iOS / Android

    The Bibliothèque Braille Romande's audiobook apps for iOS and Android started from an open-source codebase by the BSR in Lausanne. The code was refactored and the features its members asked for were added, following W3C accessibility guidelines and working with blind and partially sighted readers throughout.

    After release, the apps received a distinction from the Swiss Federal Office of Culture.

    Starting point
    An open-source codebase
    The work
    Refactoring and new features
    Recognised by
    The Swiss Federal Office of Culture
    Two iPhone screens of the BBR Player audiobook app

Book a free 15‑minute call

Pick a time that suits you: the invite reaches your inbox right away. If the time no longer works on our side, you hear from us the same working day.

What happens in 15 minutes

  1. 0–5 min

    Your situation

    What AI should do for you, and what it must never see.

  2. 5–10 min

    What's possible

    Which tasks can run privately, roughly what it costs, and what should stay human.

  3. 10–15 min

    The next step

    A workshop, a test on your own documents, or nothing yet.

  1. Pick a time
  2. Where to send the invite
Day
Time (Swiss time)
Morning
Afternoon
Prefer to write? Send a message

Tell us what you'd like AI to do, and what it must never see. We'll tell you honestly what's possible, what it costs and what should stay human. No details yet? A hello is enough.

What should we take off your plate?