PercentEncoding Class
Definition
- Assembly
- Bodu.Text.Encoding.dll
- Package
- Bodu.Text.Encoding 1.0.0
Provides span-first percent-encoding and decoding of URI components and form-style values (RFC 3986 §2.1 with the
WHATWG application/x-www-form-urlencoded rules), selecting the unescaped character set by component mode.
public static class PercentEncoding
- Inheritance
-
PercentEncoding
- Inherited Members
Examples
// URI component (default) - only unreserved characters pass through.
string component = PercentEncoding.EncodeString("a/b?c=d"); // a%2Fb%3Fc%3Dd
// Form field - space becomes '+'.
string field = PercentEncoding.EncodeString("a b+c", mode: PercentEncodingMode.FormUrlEncoded); // a+b%2Bc
// Round-trip.
string value = PercentEncoding.DecodeString("a%2Fb", mode: PercentEncodingMode.UriComponent); // a/b
Remarks
Percent-encoding represents an octet as %HH using uppercase hexadecimal digits. Each
PercentEncodingMode selects which bytes pass through unescaped - from the conservative
unreserved-only UriComponent set through to the WHATWG
FormUrlEncoded set, where space becomes +. Encoding always emits uppercase
hex; decoding accepts both uppercase and lowercase, which RFC 3986 defines as equivalent.
The byte-oriented surfaces operate on ASCII percent-encoded text; the EncodeString(string, Encoding?, PercentEncodingMode) and
DecodeString(ReadOnlySpan<char>, Encoding?, PercentEncodingMode, PercentDecodingOptions) helpers bridge .NET strings through a configurable text encoding (UTF-8 by default).
This type is not a URL parser: it does not resolve relative URLs, normalise hosts, paths, dot-segments, or scheme
casing, implement IDNA or IRI processing, accept the obsolete %uXXXX syntax, or depend on System.Web.
Because correct use depends on the component mode and decoding options, this is a static type and is intentionally not registered in BinaryEncodings - the parameterless IBinaryEncoding contract cannot carry that information.
Methods
Decode(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions)
Decodes percent-encoded characters into a byte array using the supplied mode and options.
public static byte[] Decode(ReadOnlySpan<char> source, PercentEncodingMode mode = PercentEncodingMode.UriComponent, PercentDecodingOptions options = PercentDecodingOptions.None)
Parameters
sourceReadOnlySpan<char>The percent-encoded input.
modePercentEncodingModeThe component mode.
optionsPercentDecodingOptionsThe decoding options.
Returns
- byte[]
The decoded byte array. Returns an empty array for empty input.
Exceptions
- ArgumentOutOfRangeException
Thrown when
modeis undefined.- FormatException
Thrown when the input contains a non-ASCII character or a malformed percent sequence (and AllowInvalidPercentLiterals is not set).
DecodeString(ReadOnlySpan<char>, Encoding?, PercentEncodingMode, PercentDecodingOptions)
Decodes percent-encoded characters to bytes, then converts those bytes to a string with
textEncoding (UTF-8 by default).
public static string DecodeString(ReadOnlySpan<char> source, Encoding? textEncoding = null, PercentEncodingMode mode = PercentEncodingMode.UriComponent, PercentDecodingOptions options = PercentDecodingOptions.None)
Parameters
sourceReadOnlySpan<char>The percent-encoded input.
textEncodingEncodingThe text encoding used to build the string, or null for UTF-8.
modePercentEncodingModeThe component mode.
optionsPercentDecodingOptionsThe decoding options.
Returns
- string
The decoded string.
Remarks
The default UTF8 uses a replacement decoder fallback, so invalid
UTF-8 byte sequences become the U+FFFD replacement character rather than failing. Pass a throwing encoding (for
example new UTF8Encoding(false, throwOnInvalidBytes: true)) to reject invalid sequences instead.
Exceptions
- ArgumentOutOfRangeException
Thrown when
modeis undefined.- FormatException
Thrown when the percent-encoded input is malformed.
- DecoderFallbackException
Thrown when the decoded bytes are invalid for
textEncodingand it uses a throwing fallback.
Encode(ReadOnlySpan<byte>, PercentEncodingMode)
Encodes source into a percent-encoded string using the supplied component mode.
public static string Encode(ReadOnlySpan<byte> source, PercentEncodingMode mode = PercentEncodingMode.UriComponent)
Parameters
sourceReadOnlySpan<byte>The bytes to encode.
modePercentEncodingModeThe component mode.
Returns
- string
A percent-encoded string.
Exceptions
- ArgumentOutOfRangeException
Thrown when
modeis undefined.
EncodeString(string, Encoding?, PercentEncodingMode)
Encodes a .NET string by first converting it to bytes with textEncoding (UTF-8 by default),
then percent-encoding those bytes.
public static string EncodeString(string value, Encoding? textEncoding = null, PercentEncodingMode mode = PercentEncodingMode.UriComponent)
Parameters
valuestringThe string to encode.
textEncodingEncodingThe text encoding used to obtain bytes, or null for UTF-8.
modePercentEncodingModeThe component mode.
Returns
- string
A percent-encoded string.
Exceptions
- ArgumentNullException
Thrown when
valueis null.- ArgumentOutOfRangeException
Thrown when
modeis undefined.
GetEncodedLength(ReadOnlySpan<byte>, PercentEncodingMode)
Returns the exact number of characters that Encode(ReadOnlySpan<byte>, PercentEncodingMode) produces for the supplied data and mode.
public static int GetEncodedLength(ReadOnlySpan<byte> source, PercentEncodingMode mode = PercentEncodingMode.UriComponent)
Parameters
sourceReadOnlySpan<byte>The input bytes.
modePercentEncodingModeThe component mode.
Returns
- int
The exact encoded character count.
Exceptions
- ArgumentOutOfRangeException
Thrown when
modeis undefined.
GetMaxDecodedLength(int)
Returns the maximum number of bytes that decoding charCount characters could produce.
public static int GetMaxDecodedLength(int charCount)
Parameters
charCountintThe input character count. Must be non-negative.
Returns
- int
The worst-case decoded byte count, equal to
charCount.
Exceptions
- ArgumentOutOfRangeException
Thrown when
charCountis negative.
GetMaxEncodedLength(int)
Returns the maximum number of characters that encoding byteCount bytes could produce.
public static int GetMaxEncodedLength(int byteCount)
Parameters
byteCountintThe input byte count. Must be non-negative.
Returns
- int
The worst-case encoded character count,
byteCount * 3.
Exceptions
- ArgumentOutOfRangeException
Thrown when
byteCountis negative.
IsValid(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions)
Indicates whether source is the canonical percent-encoded form for the supplied
mode - every literal character is one the mode leaves unescaped, and every percent sequence is well-formed.
public static bool IsValid(ReadOnlySpan<char> source, PercentEncodingMode mode = PercentEncodingMode.UriComponent, PercentDecodingOptions options = PercentDecodingOptions.None)
Parameters
sourceReadOnlySpan<char>The input characters.
modePercentEncodingModeThe component mode.
optionsPercentDecodingOptionsThe decoding options.
Returns
Remarks
This is stricter than Decode(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions),
which performs lenient recovery: a literal character the mode would have percent-encoded (for example #
in UriComponent, or a literal space) makes IsValid return
false even though Decode would still accept it. A percent-escaped octet such as
%2F is always canonical, so IsValid accepts it even when a literal / would not be allowed.
IsValid(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions) returning
true implies Decode succeeds, but not the converse.
TryDecode(ReadOnlySpan<char>, Span<byte>, out int, PercentEncodingMode, PercentDecodingOptions)
Attempts to decode percent-encoded characters into a destination byte span.
public static bool TryDecode(ReadOnlySpan<char> source, Span<byte> destination, out int bytesWritten, PercentEncodingMode mode = PercentEncodingMode.UriComponent, PercentDecodingOptions options = PercentDecodingOptions.None)
Parameters
sourceReadOnlySpan<char>The percent-encoded input.
destinationSpan<byte>The destination byte span.
bytesWrittenintWhen this method returns, contains the number of bytes written.
modePercentEncodingModeThe component mode.
optionsPercentDecodingOptionsThe decoding options.
Returns
TryEncode(ReadOnlySpan<byte>, Span<char>, out int, PercentEncodingMode)
Attempts to encode source into a destination character span.
public static bool TryEncode(ReadOnlySpan<byte> source, Span<char> destination, out int charsWritten, PercentEncodingMode mode = PercentEncodingMode.UriComponent)
Parameters
sourceReadOnlySpan<byte>The bytes to encode.
destinationSpan<char>The destination span.
charsWrittenintWhen this method returns, contains the number of characters written.
modePercentEncodingModeThe component mode.
Returns
TryGetDecodedLength(ReadOnlySpan<char>, out int, PercentEncodingMode, PercentDecodingOptions)
Attempts to determine the exact number of bytes that Decode(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions) will write for the supplied input.
public static bool TryGetDecodedLength(ReadOnlySpan<char> source, out int byteCount, PercentEncodingMode mode = PercentEncodingMode.UriComponent, PercentDecodingOptions options = PercentDecodingOptions.None)
Parameters
sourceReadOnlySpan<char>The input characters.
byteCountintWhen this method returns, contains the decoded byte count, or
0on failure.modePercentEncodingModeThe component mode.
optionsPercentDecodingOptionsThe decoding options.
Returns
Remarks
This mirrors Decode(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions) (lenient
recovery), not IsValid(ReadOnlySpan<char>, PercentEncodingMode, PercentDecodingOptions)
(canonical conformance): it returns the exact length Decode would write, so the result can size a decode
buffer even for decodable-but-non-canonical input.
Applies to
| Product | Versions |
|---|---|
| .NET | 8, 10 |