The Voice Match Score, explained: how we measure whether AI actually sounds like you
Most tools promise your voice and give you a vibe. Here's exactly what the Voice Match Score measures, how it's calculated, what a good score actually looks like, why we show it even when it's bad, and what happened when we checked it against an actual authorship-verification model.
Every AI writing tool claims it "sounds like you." None of them show you a number for it. That's the gap the Voice Match Score is built to close, an honest, visible measurement instead of a marketing promise.
This post explains exactly what it is, what it isn't, and why we'd rather show you a bad score than hide one.
What the score actually measures
The Voice Match Score is a 0–100 rating of how closely a generated post matches your own real writing samples. It's not a vibe check and it's not a grammar score. A generic, grammatically perfect post scores worse than a slightly rough one that matches your rhythm, because the thing being judged is similarity to you, not quality in the abstract.
Four dimensions feed the score:
- Hook pattern. How you tend to open, a claim, a question, a specific scene, a number. Compared against how the generated post opens.
- Sentence rhythm. Average sentence length and variance. Most people have a consistent rhythm they don't consciously notice, some write in short bursts, others in long connected clauses.
- Vocabulary fingerprint. The words and phrases you reach for, and (just as informative) the ones you never use. "Leverage," "delve," "game-changing" are common AI tells; if they don't appear anywhere in your samples, their appearance in a draft actively hurts the score.
- Formatting habits. Line-break density, whether you use em-dashes or semicolons, list usage, how you close a post.
How it's calculated, and what that honestly means
The score comes from a language model comparing the draft against your samples on the fixed rubric above, not from a statistical distance metric. We think it's worth saying that plainly, because the distinction matters and almost nobody in this category will tell you which one they're doing.
A model judging text on a rubric is good at what the rubric names, and blind to what it doesn't. There's published work showing this class of judge can rate a personalised draft as more like an author than that author's own writing, which tells you it's partly measuring how well the draft follows style instructions, not purely how much it reads like you. So treat the number as a reliable floor check rather than a target to optimise: a low score is strong evidence something drifted, while pushing 88 to 94 is mostly noise and not worth a credit.
That's also why we added a separate, non-AI check that flags specific phrases reading as machine-written (em-dash runs, "it's not X, it's Y," engagement bait) calibrated against your own posts so your genuine habits are never flagged. A number tells you something is off. Pointing at the sentence tells you what to fix.
We checked it against an actual authorship-verification model. Here's what we found.
"Treat it as a floor check, not a target" is an easy thing to write and a harder thing to actually verify. So we built an evaluation harness that runs our real generation pipeline against LUAR, a model built specifically to represent an author's writing style as a vector, independent of topic, and used in academic authorship-verification research rather than as a general-purpose LLM judge. It's a genuinely different kind of check than our own score: LUAR never reads the words for meaning, it just asks "does this cluster with this person's other writing, style-wise."
The setup, deliberately modeled on the same methodology as the paper cited above: four fictional authors with clearly distinct styles, each with a set of real posts held back and never shown to the generator, purely to check against afterward. For each author we generated a voice-matched post and a deliberately generic AI post from the same source material, then measured how similar each one was, in LUAR's style-space, to that author's held-back writing, and separately to a completely different author's held-back writing (the "how similar are two random people" floor).
The honest result: voice-matched generation beat the generic baseline on 3 of 4 authors, and stayed clearly above the "two random people" floor on all 4, meaning the output wasn't indistinguishable from noise. But it wasn't close to as similar to the real author as that author is to their own other writing, there's a real, measurable gap. And one author, a warm, analogy-heavy "teacher" style, scored noticeably weaker than the others; a naturally formal, businesslike voice was the one case where the generic AI post actually scored slightly closer to the real author than our voice-matched version did, plausibly because "generic AI LinkedIn voice" and "formal corporate voice" aren't that far apart to begin with, which makes them the hardest pair to tell apart.
We're publishing this because it's the same standard we're asking this whole category to be held to: a similarity claim that nobody checks against an independent, non-circular measurement is just marketing. Ours is checked now, the gap is real, and closing it (starting with seeding generation on the specific styles where it's widest) is now a tracked, numbers-backed engineering problem instead of a vibe.
What it needs to work
The score only appears when there's something real to measure against, your own posts, a website with your writing on it, or a described style with enough specificity. If you're using a default preset with no real samples of your own, we don't fabricate a score. A number with nothing behind it is worse than no number at all; it just teaches you to distrust every number a tool ever shows you.
This is also why the score is platform-specific. A voice fingerprint built from your LinkedIn posts describes your LinkedIn voice. It doesn't automatically transfer to X or Instagram, where sentence length, formality, and formatting norms are different even for the same person. Scoring an X post against LinkedIn samples would produce a number that's technically real but practically meaningless.
Reading the number
- 80–100: Close match. Minor edits at most, the structure, rhythm, and vocabulary all land inside your normal range.
- 60–79: Directionally right, noticeably generic in places. Usually fixable by adding 2–3 more real samples so the fingerprint has more to work with.
- Below 60: The generation drifted. Common causes: too few source samples (under 3), samples that aren't representative of how you actually write (e.g., marketing copy someone else wrote), or a source platform mismatch.
The score is diagnostic, not just decorative. A low score with an obvious cause (add more samples, check the platform) is more useful than a suspiciously perfect one with no explanation.
What it deliberately doesn't do
It doesn't measure whether the post is good, whether it will perform well, or whether the ideas in it are correct. A post can score 95 and still be a bad idea, badly argued. The score only tells you the voice is right. It also isn't a plagiarism or AI-detection score in reverse; it's not trying to fool a detector, it's trying to sound like a specific, real person, which happens to be the same thing that makes AI detectors quiet down, because detection systems are mostly measuring genericness, not AI use per se.
You can see it in action on the YouTube → LinkedIn tool, paste 3–5 of your real posts once, and every generation after that gets scored against them automatically.