Table of Contents

Text & Serialization

The Text & Serialization topic groups the packages that move data into and out of textual representations: binary-to-text codecs, structured document formats, and object serializers. The packages are siblings under the Bodu.Text.* prefix but do three deliberately different jobs - knowing which job you have is the entire selection problem, so this page leads with the distinction.

Three jobs that all sound like "text"

Job What it preserves Package(s)
Binary-to-text codec - bytes ⇄ printable text The exact byte sequence. No structure is added or interpreted; Encode then Decode returns the identical bytes. Bodu.Text.Encoding - Base16, Base32, Base58, Base64, Base85, plus Base45, Base62, and Bech32 / Bech32m.
Document format - parse, edit, and write structured documents The document's structure and (where the format supports it) its trivia - comments, ordering, whitespace - for faithful round-trips. Bodu.Text.Formats - Delimited (RFC 4180 CSV / TSV), DotEnv, INI; each with a typed value model and streaming readers / writers.
Object serializer - POCO ⇄ wire format Your object graph. Types, members, and collections are mapped to the format and bound back, System.Text.Json-style. Bodu.Text.Bencode (BEP 3, binary), Bodu.Text.Toml (TOML v1.0.0 / v1.1.0, text), and Bodu.Text.Yaml (YAML 1.2 core schema, text) - a shared architecture and member shape.
Text filter - select values by pattern Nothing about the values themselves - it decides only which values pass, via include/exclude glob and regex patterns compiled into one matcher. Bodu.Text.Filtering - AnyMatch sets or gitignore-style ordered rules, cost-tiered evaluation, built-in match telemetry.

The boundaries are sharp. Base64 carries any bytes but knows nothing about what they mean; an IniDocument models sections and entries but does not map them onto your types; TomlSerializer maps your types but is not a general-purpose document editor (the DOMs cover that middle ground). When two of these jobs occur together - say, a Base64-encoded blob stored inside a TOML config - you simply compose two packages.

Codecs: Bodu.Text.Encoding

Encoding families - payload expansion at a glance

The codec package fills the gaps System.Convert and System.Buffers.Text.Base64 leave open: the variants the BCL does not cover (base32hex, Crockford, z-base-32, Bitcoin / Flickr / Ripple Base58, Ascii85, Z85, Base45, Base62, Bech32 / Bech32m) and the practical surfaces around them - lenient parsing (0x prefix tolerance, whitespace stripping, missing-padding acceptance), formatting decoration (case, prefix, byte spacing, line breaks), sizing and validation helpers, and a unified IBinaryEncoding interface so the encoding can be selected at runtime from configuration.

Filters: Bodu.Text.Filtering

The filtering package selects values rather than transforming them: glob (wildcard, character-class, {a,b} alternation) and regex patterns compile once into an immutable TextFilter that classifies every pattern by evaluation cost and runs the cheapest strategies first. It speaks both established dialects - Ant / MSBuild-style include/exclude sets and gitignore-style last-match-wins ordered rules with ! negation - parses raw lines with the gitignore file conventions, and reports what matched, what was vetoed, and by which pattern through built-in statistics and an optional per-decision observer.

Document formats: Bodu.Text.Formats

The line-format libraries decode and encode self-framing documents - formats whose structure is described inline by the bytes themselves. Each of the three families (Delimited, DotEnv, INI) is a standalone System.Text.Json-shaped library: a forward-only Utf8*Reader / Utf8*Writer token pair, a *Serializer for typed binding, a mutable node DOM, and a read-only document DOM, each raising its own typed *FormatException with the source position attached.

Serializers: Bodu.Text.Bencode, Bodu.Text.Toml, and Bodu.Text.Yaml

The serializers are three libraries shaped after System.Text.Json: Bencode and TOML are member-for-member twins with only the Bencode / Toml prefix changing, and YAML shares the same architecture with a surface tuned to its format. Each layers four tiers over its format - the …Serializer for object mapping, a mutable …Node DOM for editing without a model, a read-only …Document DOM for low-allocation inspection, and the Utf8…Reader / Utf8…Writer ref-struct pair for forward-only token processing. Converters, the attribute family, naming policies, and serialization callbacks customize the mapping; those attributes, policies, and callback interfaces live in the shared Bodu.Text.Serialization package (namespace Bodu.Text.Serialization), which every serializer and line-format package references.

One nearby surface is easy to confuse with all three: the Bodu.Text namespace in Bodu.Core provides character-encoding helpers over System.Text.Encoding - BOM detection, preamble handling, span-friendly transcoding and validation. It converts bytes to characters, not bytes to a printable alphabet, and it ships in Bodu.Core, not in any package on this page. See the Bodu.Text introduction.

Note

The packages compose but do not depend on one another. Bodu.Text.Encoding and Bodu.Text.Filtering depend only on Bodu.Core; the line formats and the serializers additionally reference the small shared Bodu.Text.Serialization package for their common attribute / naming-policy / callback vocabulary. Adopting one job never pulls in the machinery of the others.

Packages in this topic

Package Status What it provides Docs
Bodu.Text.Encoding Stable Binary-to-text encodings with span / UTF-8 surfaces, OperationStatus streaming, formatting decorations, lenient parsing, and the runtime-pluggable IBinaryEncoding contract. Introduction
Bodu.Text.Filtering Preview Include/exclude text filtering: glob and regex patterns compiled into a cost-tiered TextFilter, AnyMatch sets or gitignore-style ordered rules, gitignore-convention parsing, and built-in match telemetry. Introduction
Bodu.Text.Formats Preview Self-framing document formats - Delimited (CSV / TSV), DotEnv, INI - each with a forward-only Utf8*Reader / Utf8*Writer pair, a *Serializer, and mutable / read-only DOMs. Introduction
Bodu.Text.Bencode Stable Bencode (BEP 3) serializer shaped after System.Text.Json: BencodeSerializer, mutable and read-only DOMs, and the Utf8BencodeReader / Utf8BencodeWriter ref-struct pair. Serializers introduction · Bencode
Bodu.Text.Toml Stable TOML (v1.0.0 / v1.1.0) serializer with the same member-for-member shape: TomlSerializer, both DOMs, and Utf8TomlReader / Utf8TomlWriter. Serializers introduction · TOML
Bodu.Text.Yaml Preview YAML (1.2 core schema) serializer sharing the family architecture with a YAML-tuned surface: YamlSerializer, both DOMs, the Utf8YamlReader / Utf8YamlWriter pair, block and flow collections, anchors and aliases, and multi-document streams. Serializers introduction · YAML

The authoritative dependency and status rows live in the package matrix. None of these packages depends on another package in this topic; the formats and serializers share only the Bodu.Text.Serialization vocabulary package.

Namespace orientation

The package names and root namespaces line up one-to-one, with the formats and serializers subdividing by concern:

Package Namespaces
Bodu.Text.Encoding Bodu.Text.Encoding - the per-encoding static classes, the option types, and the IBinaryEncoding registry.
Bodu.Text.Filtering Bodu.Text.Filtering - the TextFilter engine, the pattern model, options, results, and the telemetry types.
Bodu.Text.Delimited / Bodu.Text.DotEnv / Bodu.Text.Ini The standalone line-format libraries (one package per format; Bodu.Text.Formats is the umbrella meta-package over the three).
Bodu.Text.Bencode Bodu.Text.Bencode plus .Reader, .Writer, .Document, .Nodes, and .Serialization - mirroring the System.Text.Json source layout.
Bodu.Text.Toml Bodu.Text.Toml with the same .Reader / .Writer / .Document / .Nodes / .Serialization subdivision.
Bodu.Text.Yaml Bodu.Text.Yaml with the same .Reader / .Writer / .Document / .Nodes / .Serialization subdivision.

Which package do I need?

Scenario Reach for Notes
"I have bytes and need printable text" - hashes as hex, TOTP secrets, JWT segments, Bitcoin addresses, QR payloads Bodu.Text.Encoding Pick the family by expansion and alphabet - see the choose-an-encoding table.
"I select values by pattern" - keep error* lines, drop *debug* noise, honor a gitignore-style rule file Bodu.Text.Filtering Compile once, filter bulk lists cheapest-pattern-first; Evaluate reports which pattern decided - see the introduction.
"I have a CSV / .env / INI file" - parse it, walk a typed model, edit, round-trip Bodu.Text.Formats INI preserves comments and ordering on round-trip, DotEnv preserves ordering and the export flag; Delimited streams row by row.
"I map typed objects to a wire format" - config records, torrent-style payloads Bodu.Text.Toml / Bodu.Text.Bencode / Bodu.Text.Yaml Serialize / Deserialize<T> with converters, attributes, and naming policies; the three libraries share one architecture.
"I want to inspect or patch a TOML / Bencode / YAML document without a model" The serializers' DOMs Mutable …Node tree to edit, read-only …Document to inspect with minimal allocation.
"I need canonical, byte-identical output" - infohash-style hashing over the serialized form Bodu.Text.Bencode The spec mandates ascending bytewise dictionary-key order, and the serializer always emits it.
"Malformed input is expected; I don't want exceptions on the hot path" The codecs and the line formats Try* overloads and IsValid predicates on the codecs; on the line formats, the *ReaderOptions dialect policies (for example DelimitedMalformedRecordBehavior.SkipRecord) decide whether a bad record throws or is skipped.
"I need BOM detection or System.Text.Encoding helpers" The Bodu.Text namespace in Bodu.Core Character encodings, not binary-to-text codecs - see Bodu.Text and the Core Foundations topic.
"I need EditorConfig-style configuration layering over INI" Bodu.Text.Configuration Carries its own trivia-preserving INI model - see the Configuration topic.

Install

dotnet add package Bodu.Text.Encoding
dotnet add package Bodu.Text.Filtering
dotnet add package Bodu.Text.Formats
dotnet add package Bodu.Text.Toml
dotnet add package Bodu.Text.Bencode
dotnet add package Bodu.Text.Yaml

Shared design traits

However different the three jobs are, the packages share the suite's design grain, so moving between them costs little:

  • Span- and UTF-8-first. Every package exposes ReadOnlySpan<byte> / ReadOnlySpan<char> overloads alongside string and byte[]; the codecs add OperationStatus-returning streaming methods, and the serializers' readers and writers operate on UTF-8 directly.
  • Try* alongside throwing entry points. Codecs (TryDecode, TryGetDecodedLength) and validation predicates let hot paths trade exceptions for bool results; the line formats express tolerance as reader-options policies (field-count, malformed-record, duplicate-key behaviours) rather than Try* overloads.
  • Options objects, not parameter sprawl. Behavior is configured on dedicated types - BaseFormattingOptions / BaseFormatStyles for the codecs, the per-format *ReaderOptions / *WriterOptions (plus IniDocumentOptions) and *SerializerOptions for the formats, and BencodeSerializerOptions / TomlSerializerOptions / YamlSerializerOptions for the serializers.
  • Typed failures. Malformed input surfaces as a precise, format-specific exception type rather than a bare FormatException, with the failure position where the format can supply one (TOML and YAML carry line, column, and offset).
  • Streaming where the format allows it. Delimited rows, the format readers and writers, and Delimited's IAsyncEnumerable<TRecord> serializer pair process input incrementally instead of demanding the whole document in memory; the serializers' Stream overloads (sync and async) are conveniences that buffer the document in full, with only the stream copy asynchronous.

A taste of each surface

using Bodu.Text.Encoding;
using Bodu.Text.Ini;
using Bodu.Text.Toml;

// 1. Codec - bytes to printable text and back:
string hex   = Base16.Encode(hash);
byte[] token = Base64.Decode(segment, Base64Variant.UrlSafe, BaseFormatStyles.AllowMissingPadding);

// 2. Document format - parse, edit, round-trip with comments preserved:
IniObject config = IniNode.Parse(utf8Source);
config["database"].AsObject()["port"].AsValue().Value = "5433";
string updated = config.ToString();

// 3. Serializer - POCO to wire format and back:
string toml = TomlSerializer.Serialize(new AppSettings { Name = "demo", Retries = 3 });
AppSettings roundTripped = TomlSerializer.Deserialize<AppSettings>(toml);

Boundaries with neighboring packages

Three nearby surfaces sit just outside this topic, and each boundary is deliberate:

  • Bodu.Text (in Bodu.Core) - character-encoding helpers over Encoding: BOM detection, preamble handling, span-friendly transcoding and validation. Bytes to characters, not bytes to a printable alphabet. See the Bodu.Text introduction.
  • Bodu.Text.Configuration - EditorConfig-style profile presets, glob-anchored sections, and target-path resolution over its own trivia-preserving INI model, independent of Bodu.Text.Ini. When you need INI parsing alone, use Bodu.Text.Ini directly; when you need configuration layering, move up a package. See the Configuration topic.
  • Checksums and digests - the codecs print hashes; they do not compute them. Hashing lives in the Hashing & Cryptography topic.

Where to go next