Sample test report

Triage of 60 client emails, scored.

What a private AI test delivers, shown on Meow's own published test. A test on your documents follows the same form, with your documents and your answer key, under an NDA.

Task
Sort each client email for a fiduciary's team: its topics, how urgent it is, the deadline the client gives, and what the client asks for.
Documents
20 client emails written for the test, each in English, French and German: 60 in all.
Answer key
Each email labelled by hand before any model saw it, under labelling rules written down first.
Models
Open models of 0.6B, 1.7B and 8B parameters in a browser on an Apple M1 laptop; a 32B model on one NVIDIA RTX PRO 6000 at Meow's office, run a second time with the labelling rules in its instructions.
Dates
28 to 30 September 2026.

Every score

Triage, measured: size aloneCorrect answers out of 20 emails per language, with the same instructions for every model. Times are medians: 32B on a server GPU, the others on an Apple M1 laptop.
English
0.6B1.7B8B32B
Deadline19141815
Topics11121011
Urgency9111111
All three right6353
Time per emailMedian1.3 s2.4 s6.6 s1.5 s
French
0.6B1.7B8B32B
Deadline12122019
Topics1298
Urgency591412
All three right0053
Time per emailMedian1.5 s3.0 s6.7 s1.7 s
German
0.6B1.7B8B32B
Deadline13122018
Topics3746
Urgency9141010
All three right1432
Time per emailMedian1.5 s3.1 s7.4 s1.6 s
The same 32B, with the rules written inThe instructions above name the fields and nothing more. Here they also state the rules the answers were checked against, word for word, as they were written before any model saw the emails.
English
Same instructionsRules written in
Deadline1520
Topics1110
Urgency1119
All three right39
French
Same instructionsRules written in
Deadline1920
Topics86
Urgency1219
All three right36
German
Same instructionsRules written in
Deadline1820
Topics67
Urgency1020
All three right27

Every email and every answer

Where the models went wrong

  1. A bigger model closes only part of the language gap. In French, 8B found every deadline, 20 of 20, and named the topics right 9 times, against once for 0.6B. In German it also found every deadline, but filed 19 of the 20 emails under "Other".
  2. Bigger is not automatically better. With the same instructions, 32B got all three right less often than 8B in every language: it marked as urgent emails the checked answers rate as normal, and added topics the client didn't ask about.
  3. Once a model is large enough, instructions matter more than size. With the rules written in, the same 32B judged urgency right in 19 or 20 emails of 20 and found every deadline, in all three languages. Spelled-out urgency rules had made the smallest model call almost every email urgent.
  4. Topics are the hard part. Even with the rules, 32B picked exactly the right topics in only 6 to 10 emails of 20, mostly by adding a topic the email only mentions. That is what examples from your own emails are for.

The decision on these numbers

For the best run, the 32B model with the rules written in:

Part of the answerVerdictRight, out of 20
DeadlineGo20 in English, 20 in French, 20 in German
TopicsNo-go without a person checking10 in English, 6 in French, 7 in German
UrgencyGo19 in English, 19 in French, 20 in German
All three rightNo-go for full automation9 in English, 6 in French, 7 in German

The next step proposed

With the rules written in, the same 32B model went from 3 to 9 emails of 20 with all three right in English. What remains is judgement on topics: the next step is examples from your own emails, checked with the same test.

What a report on your documents adds

The cost of running it
Per document and per month, on the hardware the chosen model needs: in Infomaniak's Swiss cloud, on Meow's servers in Switzerland or at your premises.
Your documents, your answer key
30 to 50 of your real documents, read on a GPU server at Meow's office in Switzerland and scored against answers checked with you.
Written confirmation of deletion
Your documents are deleted at the end of the test, and you receive a written confirmation. The emails in this sample were written for the test, so there was nothing to delete.

Book a free 15‑minute call

Pick a time that suits you: the invite reaches your inbox right away. If the time no longer works on our side, you hear from us the same working day.

What happens in 15 minutes

  1. 0–5 min

    Your situation

    What AI should do for you, and what it must never see.

  2. 5–10 min

    What's possible

    Which tasks can run privately, roughly what it costs, and what should stay human.

  3. 10–15 min

    The next step

    A workshop, a test on your own documents, or nothing yet.

  1. Pick a time
  2. Where to send the invite
Day
Time (Swiss time)
Morning
Afternoon
Prefer to write? Send a message

Tell us what you'd like AI to do, and what it must never see. We'll tell you honestly what's possible, what it costs and what should stay human. No details yet? A hello is enough.

What should we take off your plate?