Sample test report
Triage of 60 client emails, scored.
What a private AI test delivers, shown on Meow's own published test. A test on your documents follows the same form, with your documents and your answer key, under an NDA.
- Task
- Sort each client email for a fiduciary's team: its topics, how urgent it is, the deadline the client gives, and what the client asks for.
- Documents
- 20 client emails written for the test, each in English, French and German: 60 in all.
- Answer key
- Each email labelled by hand before any model saw it, under labelling rules written down first.
- Models
- Open models of 0.6B, 1.7B and 8B parameters in a browser on an Apple M1 laptop; a 32B model on one NVIDIA RTX PRO 6000 at Meow's office, run a second time with the labelling rules in its instructions.
- Dates
- 28 to 30 September 2026.
Every score
| 0.6B | 1.7B | 8B | 32B | |
|---|---|---|---|---|
| Deadline | 19 | 14 | 18 | 15 |
| Topics | 11 | 12 | 10 | 11 |
| Urgency | 9 | 11 | 11 | 11 |
| All three right | 6 | 3 | 5 | 3 |
| Time per emailMedian | 1.3 s | 2.4 s | 6.6 s | 1.5 s |
| 0.6B | 1.7B | 8B | 32B | |
|---|---|---|---|---|
| Deadline | 12 | 12 | 20 | 19 |
| Topics | 1 | 2 | 9 | 8 |
| Urgency | 5 | 9 | 14 | 12 |
| All three right | 0 | 0 | 5 | 3 |
| Time per emailMedian | 1.5 s | 3.0 s | 6.7 s | 1.7 s |
| 0.6B | 1.7B | 8B | 32B | |
|---|---|---|---|---|
| Deadline | 13 | 12 | 20 | 18 |
| Topics | 3 | 7 | 4 | 6 |
| Urgency | 9 | 14 | 10 | 10 |
| All three right | 1 | 4 | 3 | 2 |
| Time per emailMedian | 1.5 s | 3.1 s | 7.4 s | 1.6 s |
| Same instructions | Rules written in | |
|---|---|---|
| Deadline | 15 | 20 |
| Topics | 11 | 10 |
| Urgency | 11 | 19 |
| All three right | 3 | 9 |
| Same instructions | Rules written in | |
|---|---|---|
| Deadline | 19 | 20 |
| Topics | 8 | 6 |
| Urgency | 12 | 19 |
| All three right | 3 | 6 |
| Same instructions | Rules written in | |
|---|---|---|
| Deadline | 18 | 20 |
| Topics | 6 | 7 |
| Urgency | 10 | 20 |
| All three right | 2 | 7 |
Where the models went wrong
- A bigger model closes only part of the language gap. In French, 8B found every deadline, 20 of 20, and named the topics right 9 times, against once for 0.6B. In German it also found every deadline, but filed 19 of the 20 emails under "Other".
- Bigger is not automatically better. With the same instructions, 32B got all three right less often than 8B in every language: it marked as urgent emails the checked answers rate as normal, and added topics the client didn't ask about.
- Once a model is large enough, instructions matter more than size. With the rules written in, the same 32B judged urgency right in 19 or 20 emails of 20 and found every deadline, in all three languages. Spelled-out urgency rules had made the smallest model call almost every email urgent.
- Topics are the hard part. Even with the rules, 32B picked exactly the right topics in only 6 to 10 emails of 20, mostly by adding a topic the email only mentions. That is what examples from your own emails are for.
The decision on these numbers
For the best run, the 32B model with the rules written in:
| Part of the answer | Verdict | Right, out of 20 |
|---|---|---|
| Deadline | Go | 20 in English, 20 in French, 20 in German |
| Topics | No-go without a person checking | 10 in English, 6 in French, 7 in German |
| Urgency | Go | 19 in English, 19 in French, 20 in German |
| All three right | No-go for full automation | 9 in English, 6 in French, 7 in German |
The next step proposed
With the rules written in, the same 32B model went from 3 to 9 emails of 20 with all three right in English. What remains is judgement on topics: the next step is examples from your own emails, checked with the same test.
What a report on your documents adds
- The cost of running it
- Per document and per month, on the hardware the chosen model needs: in Infomaniak's Swiss cloud, on Meow's servers in Switzerland or at your premises.
- Your documents, your answer key
- 30 to 50 of your real documents, read on a GPU server at Meow's office in Switzerland and scored against answers checked with you.
- Written confirmation of deletion
- Your documents are deleted at the end of the test, and you receive a written confirmation. The emails in this sample were written for the test, so there was nothing to delete.
Book a free 15‑minute call
Pick a time that suits you: the invite reaches your inbox right away. If the time no longer works on our side, you hear from us the same working day.
What happens in 15 minutes
- 0–5 min
Your situation
What AI should do for you, and what it must never see.
- 5–10 min
What's possible
Which tasks can run privately, roughly what it costs, and what should stay human.
- 10–15 min
The next step
A workshop, a test on your own documents, or nothing yet.
Prefer to write? Send a message
Tell us what you'd like AI to do, and what it must never see. We'll tell you honestly what's possible, what it costs and what should stay human. No details yet? A hello is enough.
