Iterate UTF-8 strings
Kinmokusei string guarantees valid UTF-8 in checked source; bstring stores arbitrary immutable bytes. Both use Go string storage. Length and indexing remain byte-oriented; range decodes code points without changing storage.
Source
import go fmt from "fmt";
function main(): void {
const text = "湯a";
fmt.Println(len(text), text[0]);
let separator = "";
for (const [offset, codePoint] of text) {
fmt.Printf("%s%d:%c", separator, offset, codePoint);
separator = " ";
}
fmt.Println();
}Run
keika check unicode-strings.km
keika run unicode-strings.kmExpected output:
4 230
0:湯 3:aRead the output
len(text)is4:湯occupies three UTF-8 bytes andaoccupies one.text[0]has typebyteand returns the first encoded byte,230, rather than a one-character string.- The two-binding range form yields an
intbyte offset and anint32Unicode code point. - The second code point begins at byte offset
3, so the offsets are0and3, not character indexes0and1.
Use bstring indexing/slicing for encoded byte protocols; string slicing panics if a boundary splits a code point. Use range when processing Unicode code points. Grapheme clusters can still contain multiple code points; user-perceived text segmentation belongs in an appropriate Go package.
Conversion and validation
This example demonstrates copied byte/code-point conversions, byte slicing, invalid UTF-8, embedded NUL and the difference between formatting and encoding:
import go { Println } from "fmt"
import go { ValidString, RuneCountInString } from "unicode/utf8"
import go { Itoa } from "strconv"
alias Bytes = byte[]
alias CodePoints = int32[]
function show(): Result<void> {
const text = "湯a"
const encoded = Bytes(text)
const decoded = CodePoints(text)
Println(len(text), len(encoded), len(decoded), string(decoded))
const snapshot = string(encoded)?
encoded[3] = 98
const changed = string(encoded)?
Println(text, snapshot, changed)
const prefix = bstring(text)[:1]
Println(len(prefix), ValidString(prefix), ValidString(text[:3]))
const raw = b"\xff\x00"
Println(len(raw), ValidString(raw), raw[1])
for (const [offset, codePoint] of raw) {
Println(offset, codePoint)
}
const [_, invalid] = string(raw)
Println(invalid !== nil)
Println(string(int32(65)), Itoa(65), string(int32(-1)) === "\ufffd")
const composed = "\u00e9"
const decomposed = "e\u0301"
Println(composed === decomposed, RuneCountInString(composed), RuneCountInString(decomposed))
return ok()
}
const main = () => {
const [err] = show()
if (err !== nil) { Println(err) }
}Expected output:
4 4 2 湯a
湯a 湯a 湯b
1 false true
2 false 0
0 65533
1 0
true
A 65 true
false 1 2The first two lines show that byte length is 4, decoded code-point count is 2, and changing a converted slice leaves both original strings unchanged. The one-byte raw prefix of 湯 is invalid UTF-8; its three-byte text prefix is valid. b"\xff\x00" retains both raw bytes: iteration reports U+FFFD (65533) for the first and NUL (0) for the second, at byte offsets 0 and 1. The following true confirms checked decoding rejects the invalid bytes.
string(int32(65)) encodes A; Itoa(65) formats the decimal string 65. Finally, the two visually similar accented strings have different bytes and code-point counts. Neither equality nor iteration normalizes text automatically.
See Types and values, Control flow, and the type-system reference.