Table of Contents

EncodingDetection Class

Definition

Namespace
Bodu.Text
Assembly
Bodu.Core.dll
Package
Bodu.Core 1.0.1
Source
EncodingDetection.cs

Provides byte-order-mark (BOM) based heuristics for detecting which Encoding was used to produce a byte sequence.

public static class EncodingDetection
Inheritance
EncodingDetection
Inherited Members

Examples

// Read a file's leading bytes and decode using the detected encoding, falling back to UTF-8.
byte[] bytes = File.ReadAllBytes(path);

System.Text.Encoding encoding =
    EncodingDetection.TryDetectByPreamble(bytes, out System.Text.Encoding? detected)
        ? detected
        : System.Text.Encoding.UTF8;

string text = encoding.GetStringSkippingPreamble(bytes);

Remarks

Detection is intentionally limited to the five canonical Unicode byte-order-marks defined by the Unicode standard:

  • UTF-8: EF BB BF
  • UTF-16 little endian: FF FE
  • UTF-16 big endian: FE FF
  • UTF-32 little endian: FF FE 00 00
  • UTF-32 big endian: 00 00 FE FF

The UTF-32 little-endian preamble shares its first two bytes with UTF-16 little endian; this implementation disambiguates them by inspecting the third and fourth bytes.

Methods

TryDetectByPreamble(ReadOnlySpan<byte>, out Encoding?)

Attempts to detect the encoding of bytes by examining its leading byte-order-mark.

public static bool TryDetectByPreamble(ReadOnlySpan<byte> bytes, out Encoding? encoding)

Parameters

bytes ReadOnlySpan<byte>

The byte span to inspect.

encoding Encoding

When this method returns true, contains the detected Encoding instance (one of UTF8, Unicode, BigEndianUnicode, or UTF32, or the UTF-32 big-endian instance returned by GetEncoding(int) for codepage 12001); otherwise null.

Returns

bool

true when a recognized BOM is detected; false when no BOM is present or when the leading bytes do not match a known preamble.

Remarks

This method does not allocate. UTF-32 little-endian detection requires inspecting four bytes; when the span is shorter than four bytes the UTF-16 little-endian preamble is reported instead, matching StreamReader behaviour.

The UTF-16 little-endian BOM (FF FE) is a strict prefix of the UTF-32 little-endian BOM (FF FE 00 00), so the two are inherently ambiguous: a UTF-16 little-endian document whose first character is U+0000 begins with the same four bytes. This method resolves the ambiguity in favour of UTF-32 little-endian whenever all four bytes match, again matching StreamReader.

Applies to

ProductVersions
.NET8, 10