The database limit is measured in bytes, while the UI shows characters.
A profile service accepts display names from browsers and stores them in a legacy field
with a four-byte UTF-8 limit. The name café has four Unicode code points, but
its UTF-8 encoding is five bytes: c, a, and f each
use one byte; é uses two. A function that slices the first four bytes can split
that last encoded character.
This example gives the storage boundary a clear contract: accept only a valid UTF-8 prefix that fits within four bytes. That does not automatically define a good product rule for names. The limit may be a legacy constraint, and a user-facing character limit should be defined separately from storage. The sample data here is illustrative; measure the actual field and the exact serialization path before deciding what failed.
- Input
café, four Unicode code points.- Storage contract
- At most 4 UTF-8 bytes, valid encoding required.
- Observed clue
- ASCII names pass; some accented names fail or appear incomplete.
- Question
- Which operation can split the encoding, and what should the diagnostic log record?
“Length” can answer several different questions.
A byte is an 8-bit storage unit. A Unicode code point is a number assigned to a character or control in the Unicode code space. An encoding such as UTF-8 maps code points to bytes. A grapheme cluster is an approximation of one user-perceived character; it may contain multiple code points, such as a base letter followed by a combining accent or an emoji joined with a modifier.
For the simple text café, there are 4 code points and 5 UTF-8 bytes. In
JavaScript, text.length counts UTF-16 code units, which can differ again: an
emoji outside the Basic Multilingual Plane takes two UTF-16 code units but one code point.
In Go, len(text) counts bytes; utf8.RuneCountInString(text) counts
decoded code points. Neither code-point count is a general grapheme-cluster count.
So write the question before reaching for a length function. A protocol frame asks for bytes. A code-point iteration asks for Unicode scalar values. A UI character counter may need grapheme segmentation. The same string can correctly have different lengths under those contracts.
| Text | UTF-8 bytes | Code points | Grapheme clusters | Why counts differ |
|---|---|---|---|---|
cafe | 4 | 4 | 4 | Four ASCII code points, one byte each. |
café | 5 | 4 | 4 | The precomposed é uses two UTF-8 bytes. |
é | 3 | 2 | 1 | e plus a combining acute accent. |
🙂 | 4 | 1 | 1 | One code point encoded as four UTF-8 bytes. |
UTF-8 uses one to four bytes for a Unicode code point.
UTF-8 represents code points with byte sequences of different lengths. ASCII values in the range U+0000–U+007F retain their familiar single-byte representation. Other values use multiple bytes, up to four. This variable length makes ordinary English-heavy text compact while allowing the same encoding to represent the full Unicode range.
The number of bytes is not a visual-width measure. A four-byte emoji may render narrow or wide depending on font and platform. A grapheme cluster may contain several code points, and a text editor can display that sequence as one user-perceived unit. Storage, cursor movement, deletion, display width, and user-facing “characters” are different operations and can need different APIs.
For café, UTF-8 bytes can be inspected as hexadecimal: 63 61 66 C3 A9. The limit of four ends between C3 and A9. A decoder cannot interpret that incomplete prefix as the original é. A safe storage boundary either rejects the whole value or truncates at a
valid encoding boundary according to an explicit product rule.
Four code points as read by a person.
Five bytes in UTF-8.
Boundary lands inside the two-byte encoding of é.
Three bytes; the next complete code point needs two.
Locate the transformation where the first bad byte appears.
When text arrives corrupted, do not assume the database caused it. Trace the value at each boundary: request decoding, application validation, normalization, serialization, database driver conversion, column storage, query decoding, and UI rendering. Compare safe measurements at each stage: declared encoding, byte length, validity, and whether the text changed. A value can be valid when received and become invalid only after a byte slice, or remain intact in storage while a UI counter misreports its length.
Use paired probes that differ in one property: cafe versus café tests multibyte encoding; precomposed é versus e plus combining
accent tests normalization; 🙂 tests both UTF-8 byte length and JavaScript UTF-16
code-unit behavior. Keep these test fixtures explicit so a future refactor does not silently
change the assumption under test.
Ask what the field is supposed to store. If it stores display text, a user-facing limit should usually count grapheme clusters and storage should still allow enough encoded bytes. If it stores protocol bytes, character-aware normalization may alter the payload and is likely inappropriate. If it stores an identifier, the accepted normalization and comparison rules belong in its specification.
Measure UTF-8 bytes and preserve complete code points.
The TypeScript example uses the browser's TextEncoder to count UTF-8 bytes,
then iterates by Unicode code point while keeping a prefix under the byte budget. The Go
example validates the string as UTF-8, uses len for bytes and utf8.RuneCountInString for code points, then chooses a rune-aligned prefix.
Both implementations are deliberate but incomplete for a user-facing character limit:
neither guarantees the result ends at a grapheme-cluster boundary. TypeScript strings can
contain unpaired UTF-16 surrogates; this example rejects them before encoding, since TextEncoder would otherwise encode the replacement character. The Go example rejects
invalid UTF-8 before slicing. A production contract should decide how malformed input is handled.
Both examples count UTF-8 bytes and return only complete encoded code points within the budget.
const encoder = new TextEncoder();
export function utf8ByteLength(text: string): number {
return encoder.encode(text).byteLength;
}
/** Keep complete Unicode code points under a UTF-8 byte budget.
* This does not preserve whole grapheme clusters such as emoji plus modifiers. */
export function prefixWithinUtf8Bytes(text: string, maxBytes: number): string {
if (!Number.isSafeInteger(maxBytes) || maxBytes < 0) {
throw new Error('maxBytes must be a non-negative safe integer');
}
let prefix = '';
let usedBytes = 0;
for (const codePoint of text) {
const firstUnit = codePoint.charCodeAt(0);
if (codePoint.length === 1 && firstUnit >= 0xd800 && firstUnit <= 0xdfff) {
throw new Error('text contains an unpaired UTF-16 surrogate');
}
const codePointBytes = utf8ByteLength(codePoint);
if (usedBytes + codePointBytes > maxBytes) break;
prefix += codePoint;
usedBytes += codePointBytes;
}
return prefix;
}
const name = 'café';
console.log({ text: name, bytes: utf8ByteLength(name), codePoints: [...name].length });
// { text: 'café', bytes: 5, codePoints: 4 }
console.log(prefixWithinUtf8Bytes(name, 4));
// caf — the next complete code point, é, needs two UTF-8 bytes.
package main
import (
"fmt"
"unicode/utf8"
)
// PrefixByBytes keeps complete UTF-8 encoded runes within a byte budget.
// It counts Unicode code points, not user-perceived grapheme clusters.
func PrefixByBytes(text string, maxBytes int) (string, error) {
if maxBytes < 0 {
return "", fmt.Errorf("maxBytes must be non-negative")
}
if !utf8.ValidString(text) {
return "", fmt.Errorf("input is not valid UTF-8")
}
used := 0
for _, r := range text {
runeBytes := utf8.RuneLen(r)
if used+runeBytes > maxBytes {
break
}
used += runeBytes
}
return text[:used], nil
}
func main() {
name := "café"
fmt.Printf("text=%q bytes=%d code-points=%d\n", name, len(name), utf8.RuneCountInString(name))
// text="café" bytes=5 code-points=4
prefix, err := PrefixByBytes("café", 4)
if err != nil {
panic(err)
}
fmt.Printf("4-byte prefix=%q bytes=%d\n", prefix, len(prefix))
// 4-byte prefix="caf" bytes=3; the next é needs two bytes.
}
Pick the unit that matches the user promise and the system boundary.
- Task A
- Explain why twelve code points and twelve grapheme clusters may not be equivalent.
- Task B
- Decide how to enforce the 48-byte storage boundary without creating invalid UTF-8.
- Transfer
- What changes if the same field is a signed protocol token rather than display text?
Show a worked answer
A product promise about visible characters should normally be implemented using a Unicode grapheme segmentation algorithm, because one displayed unit can contain several code points. A storage limit is still bytes after encoding. Validate or segment the text according to the product rule, encode it as UTF-8, measure the encoded bytes, and reject or safely truncate if it exceeds 48 bytes. Do not silently assume each grapheme costs one byte.
For an opaque signed protocol token, preserve the exact byte sequence covered by the signature. Do not normalize or truncate it as display text. Validate against the protocol's byte limit and reject an oversized token unless the protocol defines a different operation.
Further reading: Unicode FAQ on UTF encodings, Unicode FAQ on characters and combining marks, and the Go unicode/utf8 package documentation. See also the ECMAScript specification for String length and WHATWG's TextEncoder interface.