Table of Contents

Streams and token-level I/O

For inputs too large (or too hot) to materialize, each format exposes its forward-only Utf8*Reader / Utf8*Writer pair, and Delimited adds a typed record-streaming surface on its serializer.

The token surface

The readers are ref struct cursors over ReadOnlySpan<byte>: Read() advances, TokenType reports the token, GetString() decodes it, and LineNumber / BytesConsumed locate it. The writers emit UTF-8 to an IBufferWriter<byte> or a Stream (call Flush() to commit in stream mode).

Pattern 1 - walk delimited records one at a time

using Bodu.Text.Delimited;
using Bodu.Text.Delimited.Reader;

var reader = new Utf8DelimitedReader(csvBytes);
var fields = new List<string>();

while (reader.Read())
{
    switch (reader.TokenType)
    {
        case DelimitedTokenType.StartObject:   // StartArray in NoHeader mode
            fields.Clear();
            break;
        case DelimitedTokenType.String:
            fields.Add(reader.GetString());
            break;
        case DelimitedTokenType.EndObject:
            ProcessRow(fields);                // one record in memory at a time
            break;
    }
}

reader.Headers exposes the header row once it has been read.

Pattern 2 - stream typed records

For typed rows, skip the token loop:

await foreach (Trade trade in DelimitedSerializer.DeserializeAsyncEnumerableAsync<Trade>(stream))
{
    Process(trade);
}

await DelimitedSerializer.SerializeAsync(output, ProduceTradesAsync());  // IAsyncEnumerable<Trade> in

Pattern 3 - scan a DotEnv source

using Bodu.Text.DotEnv;
using Bodu.Text.DotEnv.Reader;

var reader = new Utf8DotEnvReader(envBytes);
string? key = null;

while (reader.Read())
{
    if (reader.TokenType == DotEnvTokenType.PropertyName)
        key = reader.GetString();
    else if (reader.TokenType == DotEnvTokenType.String)
        Inspect(key!, reader.GetString(), reader.LineNumber);
}

Pattern 4 - stream INI tokens as authored

using Bodu.Text.Ini;
using Bodu.Text.Ini.Reader;

var reader = new Utf8IniReader(iniBytes);

while (reader.Read())
{
    switch (reader.TokenType)
    {
        case IniTokenType.SectionHeader: EnterSection(reader.GetString()); break;
        case IniTokenType.PropertyName:  currentKey = reader.GetString(); break;
        case IniTokenType.String:        OnEntry(currentKey, reader.GetString()); break;
        case IniTokenType.Comment:       /* trivia */ break;
    }
}

Use the normalized IniDocumentReader when you want the logical object shape (globals hoisted, duplicate sections merged) instead of the physical file order - note it parses the whole document in its constructor, because merge is out-of-order.

Pattern 5 - write tokens progressively

using System.Buffers;
using Bodu.Text.Delimited.Writer;

var buffer = new ArrayBufferWriter<byte>();
var writer = new Utf8DelimitedWriter(buffer);
writer.WriteStartArray();
foreach (var row in rows)
{
    writer.WriteStartObject();
    writer.WritePropertyName("symbol"); writer.WriteString(row.Symbol);
    writer.WriteEndObject();
}
writer.WriteEndArray();
writer.Flush();

The DotEnv and INI writers are line-oriented (WritePropertyName + WriteString per entry; WriteSectionHeader / WriteComment for INI), so output is emitted as you go.

Async facades

The *Serializer stream overloads (SerializeAsync / DeserializeAsync) buffer the document in full - only the stream copy is asynchronous. The exception is Delimited's record streaming (Pattern 2), which is genuinely incremental in both directions: DeserializeAsyncEnumerableAsync reads the stream in segments and yields each record as soon as its terminating line ending is observed (memory is bounded by the longest record, not the document - a record split across segments, even inside a quoted field, is retried as more data arrives), and the IAsyncEnumerable SerializeAsync overload encodes each record as it is produced and flushes to the destination in bounded batches.

Mid-stream errors

Readers throw their *FormatException at the offending token with LineNumber / byte offset attached; everything already consumed remains valid. The ref-struct readers hold no unmanaged resources - abandoning one is safe.

See also