← All posts
Passwords

Why password strength meters lie

The green bar is counting character classes, not guesses. Here is the gap between what a meter measures and what an attacker does, why length beats symbols, and why the same password is centuries under one attack and minutes under another.

Type P@ssw0rd! into almost any sign-up form and the bar goes green. Nine characters, upper and lower case, a digit, a symbol. It satisfies every rule on the page.

It also costs an attacker with a stolen password database somewhere in the region of thirty million guesses, which a single graphics card gets through in well under a second. The meter is not broken. It is answering a different question from the one you asked it, and it never told you which.

I build login flows, and this is the thing I have had to explain most often to people who did nothing wrong: they followed the instructions on the screen, and the instructions were measuring the wrong variable.

What the green bar is actually counting

Nearly every strength meter you have met scores what Dan Wheeler, in the paper that introduced zxcvbn, named LUDS: counts of Lowercase, Uppercase, Digits and Symbols. Length gets a threshold or two, the four classes get a point each, and the total drives a colour.

That rule has one property that explains its thirty-year survival: it is four lines of code and it never needs updating. It has one other property, which is that it has almost no relationship to how hard a password is to guess.

Under LUDS, these two score identically:

PasswordLengthClassesWhat LUDS says
P@ssw0rd!94Strong
Kx7#mQvZ!94Strong

The first is the single most predictable string in this article. The second is random. A rule that cannot tell them apart is not measuring strength; it is measuring whether you read the instructions.

The attacker is not sampling the keyspace

The arithmetic behind LUDS assumes the attacker picks characters at random from the set you used. Ninety-five printable characters, nine positions, so 95^9 possibilities — about 10^17, a genuinely large number.

Nobody does this. The arithmetic describes a search nobody runs.

What actually happens is that someone loads a list. The list starts with the passwords people really choose, in order of how often they choose them, and password is in the top five of every such list ever published. Then rules are applied to every entry on the list: capitalise the first letter, append a digit, append a year, swap a for @, swap o for 0, swap e for 3, do several of those at once. Hashcat ships with rule sets that do thousands of these mutations per candidate, and they run essentially for free relative to the hashing.

So P@ssw0rd! is not one in 10^17. It is password — a top-five entry — with three mutations an attacker applies to every entry anyway, and one character on the end. The mechanism that makes it weak is not that it contains a dictionary word. It is that every transformation you applied to hide the dictionary word is itself a named, automated rule, and the rule costs the attacker a multiplier of a few, not a restart.

The trust boundary being crossed here is between you and the person who wrote the password rules. You assumed the rules encoded what makes a password hard. They encoded what is easy to check in a regular expression.

The number is missing its unit

Here is the part even the good meters get wrong, and the reason the checker we built makes it a control you can move rather than an assumption buried in the code.

A guess count is not a time. To turn guesses into a time you need a rate, and the rate depends entirely on how the service you gave the password to decided to store it — a decision you were not party to and will not learn until afterwards.

Take Summer2024!, the shape of an enormous number of corporate passwords:

The attacker hasRateSummer2024! holds for
A throttled login form100 guesses an hourover a thousand years
An API with no rate limit1,000 a seconda couple of weeks
A stolen bcrypt hash, cost 12~10,000 a secondabout a day
A stolen MD5 or SHA-1 hash~10^12 a secondunder a second

Same password. Fourteen orders of magnitude between the top row and the bottom. Every meter that shows you one bar has silently picked a row for you, and most of them pick the flattering one.

The top row is not something you earn. It is something the service grants, and it evaporates the moment their database leaks — which is the scenario the bottom row describes. Published hashcat benchmarks put a current consumer GPU in the region of 10^11 MD5 hashes a second, and renting eight of them for an evening is not an exotic capability. So the honest default is the bottom row, and that is what our checker opens on.

Length beats symbols, and it is not close

The reason is straightforward once guesses are the unit. A symbol adds roughly five bits to a random password. A word from a list of a few thousand adds eleven or twelve, and you can remember four of them.

Randall Munroe made this argument in xkcd 936 in 2011 and it has been misread ever since. The strip is not saying "use four words instead of symbols because words are magic". It is saying that the entropy sits in how the thing was chosen, not in what it is made of, and that humans are much better at remembering a long low-density secret than a short high-density one.

Run both through the same estimator against a stolen fast hash:

  • Kx7#mQ — six characters, genuinely random, all four character classes, the thing every password policy is asking for: under a second.
  • correcthorsebatterystaple — twenty-five characters, no capital, no digit, no symbol, rejected by most corporate policies: tens of thousands of years.

The second one fails the rules and beats the first by a factor you need an exponent to write down.

One honest caveat, because this is where the argument usually gets oversold: length on its own buys nothing. Twenty identical letters is twenty characters and about five hundred guesses. What length buys you is room for unpredictability — it is necessary, not sufficient, and the words have to be chosen by something that does not have taste. Four words you picked because they felt random to you are not four words of entropy.

What an honest meter would have to do

Four things, none of them expensive:

  1. Score guesses, not classes. Match the password against common-password lists, dictionaries, names, keyboard walks, dates, repeats and l33t substitutions, find the cheapest decomposition, and multiply. This is what zxcvbn does, and Wheeler showed it tracks real guessing attacks closely at the low magnitudes that actually matter.
  2. Name the attack model. A number without a rate is not an answer.
  3. Say when the user is wrong, not just that they are. "Weak" teaches nothing. "The first eight characters are one of the five most reused passwords there are, and the substitutions are a standard rule" teaches the next password too.
  4. Stop rewarding decoration. A meter that goes green when you add ! has taught the user that ! is what security looks like.

NIST got to the same place from the policy direction. The current revision of SP 800-63B, finalised in 2025, bars composition rules outright, drops mandatory periodic rotation, requires screening against breached-password blocklists, and sets a fifteen-character minimum for passwords used as a single factor. The rules that made you type P@ssw0rd! are not merely outdated; the standard that is usually cited to justify them now forbids them.

What to actually do

If you are choosing passwords. Use a password manager and let it generate long random strings you never see — that is strictly better than anything you can invent, and it removes reuse, which is the failure that actually empties accounts. Where you must remember one, use five or six words chosen by dice or by software, not by you. Where the account supports passkeys, use those instead and the whole question goes away, because there is no longer a secret the server holds.

The cost is real and I will not pretend otherwise: a manager is a single point of failure that needs its own strong master secret and a recovery plan, and a twenty-five-character passphrase is genuinely unpleasant to type into a games console. Pick where you spend it.

And go check whether the password is already out there. Strength is irrelevant for a password that has already leaked. A leaked password has a guess count of one.

If you are building the login form. Drop the composition rules; they are now non-compliant as well as counterproductive. Raise the minimum length and raise the maximum — truncating at sixteen characters is still common and it silently destroys the only variable that matters. Store passwords with a slow, memory-hard hash (argon2id, scrypt, or bcrypt at a cost you have actually benchmarked), because that is the one decision that moves your users between the top and bottom rows of the table above. Screen new passwords against a breach corpus at registration. Rate-limit by account and by source, and log the failures.

Each of those costs something. Slow hashing costs CPU on every login, so it has to be budgeted and load-tested. Blocklist screening means shipping or calling a corpus, and it means telling some users their chosen password is unacceptable for a reason they will find annoying. Raising the minimum length raises support tickets for a month. None of them costs as much as the breach.

You can run any of this through the checker. It scores entirely in your browser — no form, no upload, nothing sent anywhere — and it shows you the decomposition rather than a colour, which is the only part of this that generalises to the next password you pick.

PasswordsAuthentication