The important question is how hard it is to land on a valid secret.
A service issues an invitation token to grant access to one workspace. It is valid for 24 hours, can be redeemed once, and is accepted by an endpoint that limits attempts. The team proposes a 20-character token drawn from a 32-symbol alphabet. We will use this as an explicit, illustrative model: each of the 32 symbols is chosen with equal probability at each position, independently of the other positions.
The alphabet might be a carefully selected set of letters and digits that avoids confusing characters in a copied URL. But the human-readable design choice is only one part of the model. A secure random source and unbiased selection must make every allowed choice equally likely. A value made with a timestamp, a counter, a user name, or a general-purpose pseudorandom generator does not inherit the estimate just because it has the same visible length.
- Purpose
- Bearer token grants one-time workspace access.
- Lifetime
- 24 hours, then the invitation expires.
- Proposed model
- 20 positions; 32 equally likely symbols per position.
- Decision
- Check the outcome count and the assumptions behind it.
Bits turn a number of equally likely possibilities into a base-two scale.
One fair binary choice has two possible outcomes and represents one bit. A uniform choice
among N outcomes represents log₂(N) bits because 2log₂(N) = N. Bits are convenient because doubling the number of
possible secrets adds one bit, regardless of how large the space already is.
For a length-L string, if each position independently chooses one of A allowed symbols with equal probability, multiplication gives N = AL. Taking base-two logs turns that product into a sum: H = log₂(AL) = L × log₂(A). This is the Shannon entropy of the
stated uniform model. In security conversation it is often used as a rough guess-space
measure; it is not a complete security score.
Every position has 32 equally likely symbols.
Choices are independent across positions.
Since 32 = 2⁵, each position contributes exactly 5 modeled bits.
The same arithmetic can be checked through the full search space: 32²⁰ = (2⁵)²⁰ = 2¹⁰⁰ possible strings. An attacker who can test q distinct candidates against a
uniformly generated token succeeds with probability q / 2¹⁰⁰, provided q ≤ 2¹⁰⁰ and the candidate order gives no information. That condition is a mathematical
model, not an estimate of real attack throughput or an endorsement of a particular bit threshold.
A bit count answers one narrow question about a secret's possible values.
Entropy belongs to the process that produces values, not their printed shape.
The formula assumes every allowed string can occur and that each is equally likely. If the
generator samples each position independently and uniformly, all AL strings meet that assumption. If it picks an index by taking a random byte modulo 32, the byte
range happens to divide evenly by 32; for other alphabet sizes, modulo reduction can give some
symbols more possible source bytes than others. A standard bounded-random API can handle this
without a hand-built rejection sampler.
Dependence also changes the count. If the first position is selected from 10 templates and the remaining positions are derived from that template, the nominal character alphabet doesn't describe the reachable distribution. Predictable seeds and leaked state can narrow the possible outputs further. Conversely, cryptographic generators are designed to make outputs unpredictable to an attacker; the language's cryptographic API and the way it is used are part of the design evidence.
A secret can be hard to guess and still be easy to steal.
Entropy estimates the uncertainty of a value before an attacker learns anything about it. It does not protect a token copied into analytics, exposed in a referrer header, reused across accounts, or accepted forever. For passwords, an offline attacker who steals hashes faces a different problem from an online attacker constrained by login throttling; storage and rate limits change the threat model, not the original generation distribution.
Explore the estimate, while keeping the uniformity assumption in view.
Set a number of symbols per position and the string length. The calculator reports only the modeled entropy for independent uniform choices; it does not inspect a real password, sample a generator, or decide whether a token is safe to deploy.
20 × log₂(32) = 100.00
Assumption: every symbol is equally likely and each position is independent.
Try 16 characters from 64 equally likely choices: 16 × log₂(64) = 16 × 6 = 96 bits. Now imagine the generator chooses one of 10 templates, then fills only three positions
uniformly from those 64 symbols. The original 16 × 6 calculation no longer follows:
repeated positions are constrained by a shared template. To estimate the true distribution,
you need the generation procedure, not just the displayed alphabet and length.
Make the assumptions visible in a small calculation.
These TypeScript and Go examples calculate the same modeled quantity and reject invalid
inputs. They do not generate secrets: use the runtime's cryptographic random facilities
for production generation, and use an unbiased bounded-integer operation when selecting
from an alphabet. Node's crypto.randomInt documents an unbiased range; Go's crypto/rand package provides cryptographically secure random values and a uniform
bounded-integer function.
Both examples calculate length × log₂(alphabet size); neither one generates a secret.
/**
* Entropy estimate for a fixed-length string whose characters are selected
* independently and uniformly from the stated alphabet.
* This models the generation process; it does not estimate human choices.
*/
export function uniformStringBits(alphabetSize: number, length: number): number {
if (!Number.isSafeInteger(alphabetSize) || alphabetSize < 2) {
throw new RangeError('alphabetSize must be a safe integer of at least 2');
}
if (!Number.isSafeInteger(length) || length < 1) {
throw new RangeError('length must be a positive safe integer');
}
return length * Math.log2(alphabetSize);
}
const alphabetSize = 32;
const length = 20;
const bits = uniformStringBits(alphabetSize, length);
console.log(`${length} symbols from ${alphabetSize} choices each: ${bits} bits`);
package main
import (
"fmt"
"math"
)
// uniformStringBits estimates a fixed-length string's entropy when each
// character is selected independently and uniformly from the stated alphabet.
// It models the generation process; it does not estimate human choices.
func uniformStringBits(alphabetSize, length int) (float64, error) {
if alphabetSize < 2 {
return 0, fmt.Errorf("alphabetSize must be at least 2")
}
if length < 1 {
return 0, fmt.Errorf("length must be positive")
}
return float64(length) * math.Log2(float64(alphabetSize)), nil
}
func main() {
const alphabetSize = 32
const length = 20
bits, err := uniformStringBits(alphabetSize, length)
if err != nil {
panic(err)
}
fmt.Printf("%d symbols from %d choices each: %.0f bits\n", length, alphabetSize, bits)
}
- NIST SP 800-63B-4, Authentication and Lifecycle Management — password strength, blocklists, rate limiting, and the difficulty of estimating entropy for user-selected passwords.
- Node.js crypto.randomInt — cryptographic random integer in a range, with modulo bias avoided.
- Go crypto/rand — cryptographically secure random source and uniform bounded integer API.
Start with what the attacker can do, then decide which calculation applies.
- Name the value. Is it a public identifier, a password chosen by a person, or a bearer secret that grants access?
- Recover the generation process. Record the source of randomness, alphabet, length, selection method, seeding, and any structural constraints.
- Check the model. Only apply
L × log₂(A)when allowed choices are independent and uniform at each position. - State the attack boundary. Distinguish online attempts with throttling from offline guesses, and consider token theft, replay, and reuse separately.
- Choose evidence and controls. Inspect implementation and logs; then review expiry, one-time use, leakage paths, and consequences of compromise.
A person proposes a memorable 20-character passphrase.
Why is 20 × log₂(32) not a valid estimate merely because the password accepts 32
symbols? Describe the generation assumption that failed, then name two pieces of evidence or
controls you would examine for online guessing and for a stolen password database. Separate
your modeled calculation from what the actual password choice tells you.