EncodingDetection Class
Definition
Provides byte-order-mark (BOM) based heuristics for detecting which Encoding was used to produce a byte sequence.
public static class EncodingDetection
- Inheritance
-
EncodingDetection
- Inherited Members
Examples
// Read a file's leading bytes and decode using the detected encoding, falling back to UTF-8.
byte[] bytes = File.ReadAllBytes(path);
System.Text.Encoding encoding =
EncodingDetection.TryDetectByPreamble(bytes, out System.Text.Encoding? detected)
? detected
: System.Text.Encoding.UTF8;
string text = encoding.GetStringSkippingPreamble(bytes);
Remarks
Detection is intentionally limited to the five canonical Unicode byte-order-marks defined by the Unicode standard:
- UTF-8:
EF BB BF - UTF-16 little endian:
FF FE - UTF-16 big endian:
FE FF - UTF-32 little endian:
FF FE 00 00 - UTF-32 big endian:
00 00 FE FF
The UTF-32 little-endian preamble shares its first two bytes with UTF-16 little endian; this implementation disambiguates them by inspecting the third and fourth bytes.
Methods
TryDetectByPreamble(ReadOnlySpan<byte>, out Encoding?)
Attempts to detect the encoding of bytes by examining its leading byte-order-mark.
public static bool TryDetectByPreamble(ReadOnlySpan<byte> bytes, out Encoding? encoding)
Parameters
bytesReadOnlySpan<byte>The byte span to inspect.
encodingEncodingWhen this method returns true, contains the detected Encoding instance (one of UTF8, Unicode, BigEndianUnicode, or UTF32, or the UTF-32 big-endian instance returned by GetEncoding(int) for codepage
12001); otherwise null.
Returns
- bool
true when a recognized BOM is detected; false when no BOM is present or when the leading bytes do not match a known preamble.
Remarks
This method does not allocate. UTF-32 little-endian detection requires inspecting four bytes; when the span is shorter than four bytes the UTF-16 little-endian preamble is reported instead, matching StreamReader behaviour.
The UTF-16 little-endian BOM (FF FE) is a strict prefix of the UTF-32 little-endian BOM (FF FE 00 00),
so the two are inherently ambiguous: a UTF-16 little-endian document whose first character is U+0000 begins with
the same four bytes. This method resolves the ambiguity in favour of UTF-32 little-endian whenever all four
bytes match, again matching StreamReader.
Applies to
| Product | Versions |
|---|---|
| .NET | 8, 10 |