Are Passphrases Still Secure in the Age of AI?
Passphrases have long been considered the most pragmatic answer to weak passwords, but the usage of AI raises a fair question: does "long and random" still hold up today?
Join the DZone community and get the full member experience.
Join For FreeThe short answer up front: Yes, but only if the passphrase is genuinely random.
Passphrases emerged as a direct response to a well-known problem with classic password policies: requirements like "at least one special character, one number, one uppercase letter" tend to produce short but complex-looking passwords ("Tr0ub4dor&3") that are hard for humans to remember but comparatively easy for machines to guess because users keep falling back on the same patterns (capital letter at the start, number and special character at the end).
The passphrase flips this principle: instead of relying on character variety packed into a few positions, it relies on length through multiple words strung together. The math of passphrases makes passphrases very safe because the number of possible combinations grows exponentially with each additional word.
You can find passphrases in many parts of your life:
- Master passwords for password managers: here it's the one passphrase you still have to remember, while every other credential is generated and stored for you.
- Disk and container encryption (e.g., VeraCrypt, LUKS, BitLocker), where an attacker could attack the encrypted drive offline with no rate limiting.
- Private SSH/PGP keys, to add an extra layer of protection to the key itself in case the key file is stolen.
- Cryptocurrency wallets in the form of seed phrases (e.g., following the BIP-39 standard), which directly represent the private key.
- Wi-Fi access (WPA2/WPA3), where long passphrases make offline attacks on the handshake harder.
- Corporate and cloud logins, often combined with a second factor (MFA), replacing classic password policies.
Let's discuss for a minute what a passphrase is. A passphrase is a string of several, usually randomly chosen, words used to protect access to accounts, encrypted files, or systems. A passphrase doesn't rely on complexity through special characters — unlike passwords — but on length. The security gain comes from the sheer number of possible combinations: a random word from a list of 7,776 entries (as used in Diceware, more on that shortly) provides about 12.9 bits of entropy. Four such words combined yield roughly 51.6 bits — more than a typical eight-character password with mixed case, numbers, and special characters achieves, while being easier to remember.
The main types include:
Diceware Passphrases
With Diceware, words are selected using real dice from a fixed word list (e.g., the EFF Long Wordlist with 7,776 entries). Every combination of numbers from five dice rolls (1–6) points to exactly one word. The process can be done entirely offline and produces demonstrably random, and therefore cryptographically solid, results.
XKCD-Style ("correct horse battery staple")
Popularized by the well-known XKCD comic: four to six random, independent words strung together, often separated by hyphens or spaces. The trick: no grammatical structure, which makes it harder to guess than a meaningful sentence.
Sentence-Based Passphrases
A complete but unusual sentence, e.g., "MyCatJugglesFiveBananasOnMondays!" Advantage: easy to remember if it's personally meaningful. Disadvantage: real sentences follow linguistic patterns and are therefore easier to guess than random word lists — attackers increasingly use language models for such attacks.
Weaknesses of Passphrases
Passphrases solved well-known problems with classic passwords in many areas, but they were not the silver bullet. They had weaknesses in the past, and they will continue to have weaknesses in the future. In the age of AI, one of those weaknesses deserves particular attention, because it really puts the "long and random" principle to the test.
False randomness in AI-generated passphrases: If you have an AI generate a passphrase, the result often looks complex but follows predictable patterns and, in practice, achieves noticeably less entropy than genuine random selection.
Let's discuss this problem in more detail. The core point is that an LLM doesn't generate text through genuine randomness, but through statistical probability. Every word is chosen because it seems most plausible in context. That's exactly the opposite of what a secure passphrase needs. Even when explicitly asked to produce "random words," a model tends to favor:
- Common, everyday words over rare ones
- Certain "interesting-sounding" combinations that are overrepresented in the training corpus
- The same words across repeated requests, because the model tends toward similar outputs for similar prompts
To a human, the result looks complex and random ("fluorescence-marmalade-quantumleap"), but it's actually drawn from a much smaller effective "word pool" than a genuine random selection would be. The true number of equally likely options is smaller than the apparent complexity suggests. An attacker who knows that a passphrase was generated by a specific AI doesn't need to search the entire character space — they can instead build a targeted dictionary from that model's preferred patterns and smaller effective word pool, and run a guessing attack against it.
Similar Effects Have Already Been Observed for AI-Generated Passwords
The problem described above isn't merely an academic exercise — early data on this topic already exists, though so far it has been studied analogously for AI-generated passwords. In February 2026, security firm Irregular published a report defining this exact concept as a viable attack method against AI-generated passwords, and Kaspersky's cracking tests confirmed its practical effectiveness. Specifically: once an attacker identifies which LLM generated a target credential, an exhaustive brute-force attack against the full 94^16 character space is no longer necessary. Using a model-specific attack dictionary, candidates can be ranked by their empirical generation frequency and searched via a probabilistically optimized attack against a key space that's smaller by several orders of magnitude.
A separate analysis published on VPN Central in February 2026 examined the same type of attack and found that, combined with brute-force techniques, credentials could be recovered within minutes.
Conclusion
For passwords, several independent studies have documented with concrete figures that AI-generated passwords carry less entropy. For passphrases specifically, the current body of data is considerably thinner. However, a plausible extension of the same underlying mechanism — LLMs optimize for probability rather than randomness, which applies to words just as much as to characters - is a legitimate inference, and it suggests that AI-generated passphrases likely carry noticeably less entropy than genuinely random ones. The study "LLM-Generated Passphrases That Are Secure and Easy to Remember" (ACL Anthology, NAACL 2025 Findings) provides indirect data on this question. It's worth noting, though, that the study is primarily concerned with defining the conditions under which an LLM can be made to produce usable passphrases in the first place.
In short: numeric, evidence-backed data exists for passwords; for passphrases, the same mechanism — statistical bias rather than genuine randomness — is plausible and architecturally well-grounded, but still thinly supported empirically.
Opinions expressed by DZone contributors are their own.
Comments