Skip to content

Iterate UTF-8 strings ​

Kinmokusei string guarantees valid UTF-8 in checked source; bstring stores arbitrary immutable bytes. Both use Go string storage. Length and indexing remain byte-oriented; range decodes code points without changing storage.

Source ​

ts
import go fmt from "fmt";

function main(): void {
  const text = "湯a";
  fmt.Println(len(text), text[0]);

  let separator = "";
  for (const [offset, codePoint] of text) {
    fmt.Printf("%s%d:%c", separator, offset, codePoint);
    separator = " ";
  }
  fmt.Println();
}

Run ​

sh
keika check unicode-strings.km
keika run unicode-strings.km

Expected output:

text
4 230
0:湯 3:a

Read the output ​

  • len(text) is 4: 湯 occupies three UTF-8 bytes and a occupies one.
  • text[0] has type byte and returns the first encoded byte, 230, rather than a one-character string.
  • The two-binding range form yields an int byte offset and an int32 Unicode code point.
  • The second code point begins at byte offset 3, so the offsets are 0 and 3, not character indexes 0 and 1.

Use bstring indexing/slicing for encoded byte protocols; string slicing panics if a boundary splits a code point. Use range when processing Unicode code points. Grapheme clusters can still contain multiple code points; user-perceived text segmentation belongs in an appropriate Go package.

Conversion and validation ​

This example demonstrates copied byte/code-point conversions, byte slicing, invalid UTF-8, embedded NUL and the difference between formatting and encoding:

ts
import go { Println } from "fmt"
import go { ValidString, RuneCountInString } from "unicode/utf8"
import go { Itoa } from "strconv"

alias Bytes = byte[]
alias CodePoints = int32[]

function show(): Result<void> {
  const text = "湯a"
  const encoded = Bytes(text)
  const decoded = CodePoints(text)
  Println(len(text), len(encoded), len(decoded), string(decoded))

  const snapshot = string(encoded)?
  encoded[3] = 98
  const changed = string(encoded)?
  Println(text, snapshot, changed)

  const prefix = bstring(text)[:1]
  Println(len(prefix), ValidString(prefix), ValidString(text[:3]))

  const raw = b"\xff\x00"
  Println(len(raw), ValidString(raw), raw[1])
  for (const [offset, codePoint] of raw) {
    Println(offset, codePoint)
  }
  const [_, invalid] = string(raw)
  Println(invalid !== nil)

  Println(string(int32(65)), Itoa(65), string(int32(-1)) === "\ufffd")
  const composed = "\u00e9"
  const decomposed = "e\u0301"
  Println(composed === decomposed, RuneCountInString(composed), RuneCountInString(decomposed))
  return ok()
}

const main = () => {
  const [err] = show()
  if (err !== nil) { Println(err) }
}

Expected output:

text
4 4 2 湯a
湯a 湯a 湯b
1 false true
2 false 0
0 65533
1 0
true
A 65 true
false 1 2

The first two lines show that byte length is 4, decoded code-point count is 2, and changing a converted slice leaves both original strings unchanged. The one-byte raw prefix of 湯 is invalid UTF-8; its three-byte text prefix is valid. b"\xff\x00" retains both raw bytes: iteration reports U+FFFD (65533) for the first and NUL (0) for the second, at byte offsets 0 and 1. The following true confirms checked decoding rejects the invalid bytes.

string(int32(65)) encodes A; Itoa(65) formats the decimal string 65. Finally, the two visually similar accented strings have different bytes and code-point counts. Neither equality nor iteration normalizes text automatically.

See Types and values, Control flow, and the type-system reference.

Kinmokusei is a pre-1.0 project. Documentation describes implemented behavior.