CatMind AI logoCatMind

How CatMind stays honest

Cats don't speak in sentences, so CatMind doesn't pretend they do. Here is exactly what it does, what it refuses to do, and what happens to your recordings.

Probabilities, never translation

CatMind never claims to translate cat language. Every reading is phrased as what the behaviour may indicate, with a confidence level and ranked alternatives you can weigh yourself.

Nothing is invented

If a signal isn't visible or audible in your recording, CatMind says so — 'not clearly visible' rather than filling the gap. When nothing fits, the interpretation is simply 'unclear'.

Your context is context, not truth

What you tell CatMind about the moment — where it happened, when your cat last ate, what you think was going on — informs the reading but never overrides what is actually observed.

Never a diagnosis

Signals that might relate to discomfort are reported as observations only, always alongside a reminder to speak to a veterinarian. CatMind is not a medical tool.

Private recordings

Your videos, audio and photos are stored privately and are only ever reachable through short-lived links tied to your account. They are used to produce your results, nothing else.

Deletable at any time

You can delete a single analysis, a recording, a whole cat with everything learned about them, or your entire account and all of its data.

How confidence is described

  • 90–100% — very strong evidence in the recording
  • 75–89% — strong evidence
  • 55–74% — moderate evidence
  • 35–54% — weak evidence, treat as a possibility
  • 0–34% — low; mentioned only for completeness

Confidence is deliberately lowered when the recording is short or unclear, when several cats appear, or when the signals point in more than one direction.

How CatMind learns your cat

Personalisation comes from remembering your cat, not from changing the underlying model. A pattern starts as still learning, becomes learned after 3 confirmations, and strong after 5 consistent ones. If your corrections contradict a pattern, it loses confidence and can be marked disputed. Only a handful of the most relevant patterns are used for any single analysis — never your whole history.

Live listening, and cues sent back

In a live session, your microphone stays open on your device only. CatMind measures the sound level on your device to notice audio activity and cuts out the short clips around it. Only those clips are analysed — continuous silence is never uploaded and never costs an analysis. The activity detector is not a cat detector: whether a sound came from a cat is decided by the analysis itself, and when there isn't enough evidence CatMind says so instead of guessing.

Cues you send are experimental signals. When you write something like “come here”, CatMind maps it to one safe communication intent and prepares a short, gentle synthetic tone. It is not translated cat speech and it is not a recording of a cat. CatMind will never generate threatening, hissing, growling, distress or mating sounds — those intents do not exist in the taxonomy at all. Cues never play by themselves; you always press Play, and you control the volume.

After a cue plays, CatMind listens for a short response window and links any sound it hears to that cue, then asks you whether your cat responded as expected. Those ratings — and only those ratings, from you — build up what CatMind knows about which cues suit your cat. No percentage is shown before at least three reviewed plays.

What is planned but not running

Specialized audio models — an AST-style audio transformer and CLAP audio embeddings — are designed into CatMind as a future external service: a cat/not-cat gate, general call and context models, and per-cat similarity search over embeddings. None of them runs today. When a screen shows a sound reading, it comes from CatMind's own measurements plus the multimodal analysis, never from a model we haven't built yet. Research datasets stay non-commercial reference and benchmark material only.

Where the sound knowledge comes from

When a recording contains sound, CatMind describes the call it hears — meow, trill, purr, yowl, hiss, growl or chatter — and compares it with three documented recording contexts from published research: waiting for food, being brushed or handled, and being isolated in an unfamiliar place. No research audio is stored in CatMind and no model is retrained; the research is used as reference knowledge only.

What we measured: on 15 September 2026 we ran 20 labelled research recordings (7 waiting for food, 7 isolation in an unfamiliar place, 6 brushing) through the app exactly as a user would. A cat vocalisation was correctly detected in 20 out of 20. The closest matching research context was the true one in 7 out of 20 — no better than chance for three contexts. Because of that result, asking the language model to judge the recording context by ear is no longer trusted on its own.

What we do instead: CatMind now measures your recording itself. Before anything is sent for interpretation, the app analyses the actual audio signal — how many separate calls there are, how long each one lasts, the gaps between them, pitch and pitch variation, brightness and loudness. Those measurements are facts about your recording, not impressions. They are then scored by a small acoustic model built inside CatMind from all 440 labelled CatMeows recordings (21 cats). Tested on cats it had never heard, it names the correct context 49% of the time, against 33% for guessing. That is real signal but still weak, so it only nudges confidence slightly, and sound-only recordings stay capped well below what video allows. No dataset audio ships with the app and no model is trained on your recordings.

Naming the sound type: CatMind also names the type of sound it measured — meow, trill, chatter, purr, murmur, yowl, caterwaul, hiss, growl or shriek. These are the ten sound types described in the Meowsic research project, and they are sound categories, not translated words. The decision is made from the measured numbers of your recording compared against the published acoustic description of each type. Because no openly licensed set of recordings labelled with these ten types is available to us, this part carries no measured accuracy figure of its own: it improves how the sound is described, and it does not raise the confidence of any interpretation. When the measurements don't clearly match one type, CatMind says "unclear" instead of naming one.

  • · CatMeows (Ludovico et al., Zenodo record 4008297) — labelled meows in three contexts, licensed for non-commercial use (CC BY-NC 4.0).
  • · Cat vocalisation recordings (Zenodo record 4724180) — call types beyond the meow.
  • · Meowsic (Susanne Schötz, Lund University) — the typology of ten cat sound types and their acoustic descriptions, used as written reference only. No Meowsic audio or video is downloaded, stored or trained on; training use would need the researcher's written permission.
  • · Battersea — published guidance on visible cat body language, used to keep CatMind's wording of body-language signals consistent.
  • · Public cat-vocalisation classification work (IsolaHGVIS/Cat-Meow-Classification, Ikteder/cat-vocalization-interpreter) — category structure and feature-to-meaning conventions.

The research sources registry documents every dataset and model behind the sound reading — what the data is, how it is licensed, its measured accuracy, and how the list updates itself when new research is published.

CatMind provides behavioural interpretations, not literal translation or veterinary diagnosis.