6 min read

Source: Roblox Creator Hub · CC BY 4.0 · View source · Code samples: MIT Imported 2026-10-03. Formatting adapted for this site.

utf8

This library provides basic support for UTF-8 encoding. This library does not provide any support for Unicode other than the handling of the encoding. Any operation that needs the meaning of a character, such as character classification, is outside its scope.

Unless stated otherwise, all functions that expect a byte position as a parameter assume that the given position is either the start of a byte sequence or one plus the length of the subject string. As in the string library, negative indices count from the end of the string.

You can find a large catalog of usable UTF-8 characters here.

Properties

NameType / ReturnsDescription
utf8.charpatternstringThe pattern "[%z\x01-\x7F\xC2-\xF4][\x80-\xBF]*", which matches exactly zero or more UTF-8 byte sequences, assuming that the subject is a valid UTF-8 string.

utf8.charpattern

The pattern "[%z\x01-\x7F\xC2-\xF4][\x80-\xBF]*", which matches exactly zero or more UTF-8 byte sequence, assuming that the subject is a valid UTF-8 string.

FieldValue
typestring

Functions

NameType / ReturnsDescription
utf8.charstringConverts zero or more codepoints to UTF-8 byte sequences.
utf8.codesfunction, string, intReturns an iterator function that iterates over all codepoints in a given string.
utf8.codepointTupleReturns the codepoints (as integers) from all codepoints in a given string.
utf8.lenintReturns the number of UTF-8 codepoints in a given string.
utf8.offsetint?Returns the position (in bytes) where the encoding of the n‑th codepoint of s (counting from byte position i) starts.
utf8.graphemesfunctionReturns an iterator function that iterates over the grapheme clusters of a given string.
utf8.nfcnormalizestringConverts the input string to Normal Form C.
utf8.nfdnormalizestringConverts the input string to Normal Form D.

utf8.char

Receives zero or more codepoints as integers, converts each one to its corresponding UTF-8 byte sequence and returns a string with the concatenation of all these sequences.

Parameters

NameTypeDefaultDescription
codepointsTupleZero or more integer codepoints in the range [0, 0x10ffff] to encode.

Returns

TypeDescription
stringA string containing the concatenated UTF-8 byte sequences for each codepoint.

utf8.codes

Returns an iterator function so that the construction:

local str = "héllo"
for position, codepoint in utf8.codes(str) do
	print(position, codepoint)
end

will iterate over all codepoints in string str. It raises an error if it meets any invalid byte sequence.

Parameters

NameTypeDefaultDescription
strstringThe string to iterate over.

Returns

TypeDescription
functionThe iterator function that, on each call, returns the byte position and codepoint of the next character.
stringThe input string (passed as the invariant state to the iterator).
intThe initial control value (0) for the iterator.

utf8.codepoint

Returns the codepoints (as integers) from all codepoints in the provided string (str) that start between byte positions i and j (both included). The default for i is 1 and for j is i. It raises an error if it meets any invalid byte sequence.

Parameters

NameTypeDefaultDescription
strstringThe UTF-8 encoded string to extract codepoints from.
iint1The index of the codepoint that should be fetched from this string.
jintiThe index of the last codepoint between i and j that will be returned. If excluded, this will default to the value of i.

Returns

TypeDescription
TupleThe codepoints as integers for all characters that start between byte positions i and j.

utf8.len

Returns the number of UTF-8 codepoints in the string str that start between positions i and j (both inclusive). The default for i is 1 and for j is -1. If it finds any invalid byte sequence, returns a nil value plus the position of the first invalid byte.

Parameters

NameTypeDefaultDescription
sstringThe UTF-8 encoded string to measure.
iint1The starting position.
jint-1The ending position.

Returns

TypeDescription
intThe number of UTF-8 codepoints in the specified range, or nil followed by the byte position of the first invalid byte sequence.

utf8.offset

Returns the position (in bytes) where the encoding of the n‑th codepoint of s (counting from byte position i) starts. A negative n gets characters before position i. The default for i is 1 when n is non-negative and #s + 1 otherwise, so that utf8.offset(s, -n) gets the offset of the n‑th character from the end of the string. If the specified character is neither in the subject nor right after its end, the function returns nil.

Parameters

NameTypeDefaultDescription
sstringThe UTF-8 encoded string to search within.
nintThe character offset to seek. Positive values count forward, negative values count backward, and 0 finds the start of the character at byte position i.
iint1The byte position from which to start counting. Defaults to 1 when n is non-negative and #s + 1 when n is negative.

Returns

TypeDescription
int?The byte position where the target codepoint begins, or nil if the character is not within the string.

utf8.graphemes

Returns an iterator function so that

for first, last in utf8.graphemes(str) do
	local grapheme = s:sub(first, last)
	-- body
end

will iterate the grapheme clusters of the string.

Parameters

NameTypeDefaultDescription
strstringThe UTF-8 encoded string to iterate over for grapheme clusters.
inumberThe starting byte position within the string. Negative values count from the end.
jnumberThe ending byte position within the string. Negative values count from the end.

Returns

TypeDescription
functionAn iterator function that returns the start and end byte positions of each grapheme cluster.

utf8.nfcnormalize

Converts the input string to Normal Form C, which tries to convert decomposed characters into composed characters.

Parameters

NameTypeDefaultDescription
strstringThe UTF-8 encoded string to normalize.

Returns

TypeDescription
stringThe NFC-normalized string with decomposed characters composed into their precomposed equivalents.

utf8.nfdnormalize

Converts the input string to Normal Form D, which tries to break up composed characters into decomposed characters.

Parameters

NameTypeDefaultDescription
strstringThe string to convert.

Returns

TypeDescription
stringThe NFD-normalized string with precomposed characters decomposed into their component parts.