mirror of
https://github.com/gravitational/teleport.git
synced 2026-10-11 22:49:54 +00:00
appresource: Add path tokenizer and validator (#68795)
* appresource: Add path tokenizer and validator Create the `lib/appresource` package with the request-path tokenizer and wire-form validator, promoted from the sketch branch. `Tokenize` validates a raw request path and splits it into segments without decoding the bytes it forwards. A separate decode is used only as a validation view. Tokenize rejects any byte outside the RFC 3986 path set less ";", admits the encoded separator `%2F`/`%2f` and percent-encoded non-ASCII content as the only escapes, rejects "." and ".." segments and "//" on a view that also decodes the encoded separator, and requires each segment, decoded once, to be valid UTF-8, NFKC-stable, and made of graphic runes. Together these prevent the canonicalization and fold bypasses where an upstream decodes or normalizes a path onto a segment the matcher never saw. Paths over 8 KiB are rejected. The package is a leaf so the app agent, role validation in `lib/services`, and tctl can all import it. Node matching, captures, and the encoded-slash grammar follow in later changes. * appresource: Validate path content per segment The NFKC and graphic-rune checks ran on the whole path body with the encoded separator `%2F` kept literal. A combining mark placed right after an encoded slash then composed with that literal `F` (`F` plus U+0307 forms `Ḟ`), so a legitimate path was rejected as not NFKC-normalized even though its real decoded content is normal. Move the content checks into `validateDecoded` so they run per segment on the decode-for-validation view, where `%2F` is a separator. An encoded slash and a raw `/` now validate the same way, and the merge drops a second `decodeSeparators` pass and split. Pin the non-breaking-space reject case to the NFKC rule, since the graphic check would otherwise let it pass for the wrong reason. * appresource: Address tokenizer review feedback Reject a segment that starts with a combining mark. The mark pairs with the previous rune, and net/http.ServeMux still splits on the literal slash, so upstreams may disagree on the segment boundary. RFC 5891 imposes the same rule on the start of a domain label. Re-add the rapid property tests. Unlike the fuzz target, which only replays its seed corpus under go test, they exercise the tokenizer invariants on every run. Register the tokenizer fuzz target in fuzz/oss-fuzz-build.sh and pin the encoded hash %23 (#) as rejected. * appresource: Allow the encoded space %20 in paths Allow %20 as content, alongside the non-ASCII escapes. A space has no raw form in a URL path, so the escape is its only representation, and rejecting it makes resources unreachable with no workaround. Jenkins serves job names with spaces at /job/My%20Job/lastBuild, and SharePoint serves its default library at Shared%20Documents. A space cannot decode into a separator, so it needs neither a rule-global grant nor a per-segment pattern, unlike the encoded separator %2F. Reject a segment that starts or ends with a space. Windows and IIS trim a trailing space from a segment, so ".. " would resolve as ".." and "secret " would resolve as "secret", escaping a carve-out. That is the same fold-to-a-different-segment problem the other content checks prevent, reached by trimming rather than normalization. Reject a segment of only dots and spaces, which folds the former "." and ".." equality check into one rule. A normalizer that strips spaces or a single trailing dot resolves ". ." and "..." the same way. Permit U+0020 alone in isGraphicRune, never the Zs category. U+1680 ogham space mark, U+2028, U+200B and U+0085 are NFKC-stable, so the graphic-rune check is the only place that rejects them. * appresource: Reject a segment's trailing dots Reject a segment whose trailing run of dots and spaces contains a dot. IIS and Windows trim trailing dots and spaces from a path component, so an upstream would resolve `secret.` and `secret%20.` as `secret`, a segment the matcher never saw. The edge-space rule already closes this trim vector for the space; close it for the dot the same way. A leading dot stays allowed, so `/.well-known` remains reachable. Fold the escape predicate shared by `validatePercentEscapes` and `decode` into `isAllowedEscape`, so the two functions visibly accept the same escape set. Clip paths and segments to 256 bytes in error messages, because a rejected path can be kilobytes long and the message survives into logs. Pin the acceptance of a combining mark riding a trailing interior space, which ends the segment with a mark rather than a space byte, so no upstream trim fires on it. * appresource: Reject a leading conjoining jamo Extend the leading-mark check with the NFKC boundary property. A conjoining Hangul jamo is a letter, not a mark, so unicode.IsMark misses it, yet it composes onto the character before it the same way. A segment of "ᆨ" after a segment ending in "가" joins to "각" once an upstream drops the separator, which is the fold the check prevents. Render an illegal byte as a one-byte slice instead of converting it to a rune, so a raw 0xC3 prints as "\xc3" rather than "Ã". * appresource: Simplify godocs and merge byte checks Drop the numbered rule list from the Tokenize godoc, along with the terms NFKC-stable, graphic rune, and conjoining Hangul jamo that a caller does not need in order to call the function correctly. Say what to pass, what comes back, and which paths are rejected, and keep the precise rules on the reject helpers next to the code that enforces them. Add an inline comment where the leading mark check needs both unicode.IsMark and the NFKC boundary property, since neither covers the other. Rename lengthCap to maxPathLength, which reads better at the call site, and merge validatePathBytes and validatePercentEscapes into validateRawBytes so the raw path is walked once rather than twice. Both loops were byte-level and neither depended on the other having finished, and the hex digits the escape branch skips are alphanumeric, so they were always legal path bytes anyway. Swap the formatted testify variants in the property tests for the plain ones, which take the same message and arguments.
This commit is contained in:
@@ -137,6 +137,9 @@ build_teleport_fuzzers() {
|
||||
compile_native_go_fuzzer $TELEPORT_PREFIX/lib/scopes \
|
||||
FuzzValidateQualifiedName fuzz_validate_qualified_name
|
||||
|
||||
compile_native_go_fuzzer $TELEPORT_PREFIX/lib/appresource \
|
||||
FuzzTokenizeNonASCII fuzz_tokenize_non_ascii
|
||||
|
||||
}
|
||||
|
||||
build_teleport_api_fuzzers() {
|
||||
|
||||
@@ -0,0 +1,266 @@
|
||||
/*
|
||||
* Teleport
|
||||
* Copyright (C) 2026 Gravitational, Inc.
|
||||
*
|
||||
* This program is free software: you can redistribute it and/or modify
|
||||
* it under the terms of the GNU Affero General Public License as published by
|
||||
* the Free Software Foundation, either version 3 of the License, or
|
||||
* (at your option) any later version.
|
||||
*
|
||||
* This program is distributed in the hope that it will be useful,
|
||||
* but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
* GNU Affero General Public License for more details.
|
||||
*
|
||||
* You should have received a copy of the GNU Affero General Public License
|
||||
* along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
*/
|
||||
|
||||
// Package appresource checks whether an HTTP app request is allowed
|
||||
// by a role. Roles carry allow-only rules. A rule can match on
|
||||
// request path, HTTP method, and a where predicate over the user
|
||||
// identity. Every field is optional.
|
||||
//
|
||||
// Example role fragment:
|
||||
//
|
||||
// allow:
|
||||
// app_resources:
|
||||
// - paths:
|
||||
// - /api/v4/user/{username}
|
||||
// where: user.name == vars.username
|
||||
package appresource
|
||||
|
||||
import (
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode"
|
||||
"unicode/utf8"
|
||||
|
||||
"github.com/gravitational/trace"
|
||||
"golang.org/x/text/unicode/norm"
|
||||
)
|
||||
|
||||
// maxPathLength bounds the length of a path Tokenize accepts.
|
||||
const maxPathLength = 8 << 10 // 8 KiB
|
||||
|
||||
// legalPathPunct is the non-alphanumeric bytes allowed in a raw
|
||||
// URL path. It is RFC 3986 pchar except for ";", plus "/" and "%".
|
||||
//
|
||||
// ";" is dropped because matrix parameters and ";jsessionid" may
|
||||
// cause the matcher and the upstream app to disagree on where the
|
||||
// path ends.
|
||||
const legalPathPunct = "-._~!$&'()*+,=:@/%"
|
||||
|
||||
// Tokenize validates an HTTP request path and splits it on a real
|
||||
// "/" into the encoded segments a role's path rules match against, so
|
||||
// an encoded slash stays inside one segment. Pass
|
||||
// [net/url.URL.EscapedPath], the encoded path sent to the upstream
|
||||
// app, not the already-decoded [net/url.URL.Path].
|
||||
//
|
||||
// Tokenize accepts a path that starts with "/", stays under 8 KiB,
|
||||
// and holds only the path characters RFC 3986 allows, except for ";".
|
||||
// Anything else has to be sent percent-encoded, and the only escapes
|
||||
// allowed are the separator %2F ("/"), the space %20 (" "), and the
|
||||
// UTF-8 bytes of non-ASCII text. Tokenize also rejects a path an
|
||||
// upstream app could read as a different path than the one a role
|
||||
// matched, such as "/a/../b" or "/files/secret." on a server that
|
||||
// trims a trailing dot.
|
||||
func Tokenize(path string) ([]string, error) {
|
||||
if len(path) > maxPathLength {
|
||||
return nil, trace.BadParameter("path length %d exceeds the %d byte limit", len(path), maxPathLength)
|
||||
}
|
||||
if !strings.HasPrefix(path, "/") {
|
||||
return nil, trace.BadParameter("path %q must start with /", clip(path))
|
||||
}
|
||||
if err := validateRawBytes(path); err != nil {
|
||||
return nil, trace.Wrap(err)
|
||||
}
|
||||
if err := validateDecoded(path); err != nil {
|
||||
return nil, trace.Wrap(err)
|
||||
}
|
||||
return strings.Split(path[1:], "/"), nil
|
||||
}
|
||||
|
||||
// validateRawBytes rejects any byte that cannot appear in a URL path
|
||||
// under RFC 3986, any invalid percent-escape, and every escape except
|
||||
// %2F ("/"), %20 (" "), or one that decodes to a non-ASCII byte.
|
||||
func validateRawBytes(path string) error {
|
||||
for i := 0; i < len(path); i++ {
|
||||
if !isLegalPathByte(path[i]) {
|
||||
return trace.BadParameter("path %q contains an illegal URL byte %q", clip(path), path[i:i+1])
|
||||
}
|
||||
if path[i] != '%' {
|
||||
continue
|
||||
}
|
||||
if i+2 >= len(path) {
|
||||
return trace.BadParameter("path %q has a truncated percent-escape", clip(path))
|
||||
}
|
||||
v, err := strconv.ParseUint(path[i+1:i+3], 16, 8)
|
||||
if err != nil {
|
||||
return trace.BadParameter("path %q has a malformed percent-escape %q", clip(path), path[i:i+3])
|
||||
}
|
||||
if !isAllowedEscape(byte(v)) {
|
||||
const msg = "path %q contains the percent-escape %q; only the encoded separator %%2F, the encoded space %%20, and non-ASCII content escapes are allowed"
|
||||
return trace.BadParameter(msg, clip(path), path[i:i+3])
|
||||
}
|
||||
i += 2
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// isLegalPathByte reports whether a given byte may appear in a raw
|
||||
// URL path.
|
||||
func isLegalPathByte(b byte) bool {
|
||||
switch {
|
||||
case b >= 'A' && b <= 'Z', b >= 'a' && b <= 'z', b >= '0' && b <= '9':
|
||||
return true
|
||||
}
|
||||
return strings.IndexByte(legalPathPunct, b) >= 0
|
||||
}
|
||||
|
||||
// isAllowedEscape reports whether a percent-escape decoding to b may
|
||||
// appear in a path. [validateRawBytes] accepts exactly the escapes
|
||||
// [decode] resolves, which are the separator, the space, and
|
||||
// non-ASCII content.
|
||||
func isAllowedEscape(b byte) bool {
|
||||
return b == '/' || b == ' ' || b >= 0x80
|
||||
}
|
||||
|
||||
// validateDecoded checks the decoded validation view of the path. It
|
||||
// rejects consecutive slashes, "." and ".." segments, and any content
|
||||
// that is not NFKC-stable graphic UTF-8.
|
||||
func validateDecoded(path string) error {
|
||||
decoded := decode(path)
|
||||
if strings.Contains(decoded, "//") {
|
||||
const msg = "path %q has consecutive slashes once the encoded separator %%2F is decoded"
|
||||
return trace.BadParameter(msg, clip(path))
|
||||
}
|
||||
if !utf8.ValidString(decoded) {
|
||||
return trace.BadParameter("path %q is not valid UTF-8 once decoded", clip(path))
|
||||
}
|
||||
for seg := range strings.SplitSeq(decoded[1:], "/") {
|
||||
if err := rejectDotSegment(seg); err != nil {
|
||||
return trace.Wrap(err)
|
||||
}
|
||||
if err := rejectLeadingMark(seg); err != nil {
|
||||
return trace.Wrap(err)
|
||||
}
|
||||
if err := rejectEdgeSpace(seg); err != nil {
|
||||
return trace.Wrap(err)
|
||||
}
|
||||
if err := rejectTrailingDot(seg); err != nil {
|
||||
return trace.Wrap(err)
|
||||
}
|
||||
}
|
||||
if !norm.NFKC.IsNormalString(decoded) {
|
||||
return trace.BadParameter("path %q is not NFKC-normalized", clip(path))
|
||||
}
|
||||
for _, r := range decoded {
|
||||
if !isGraphicRune(r) {
|
||||
const msg = "path %q contains the disallowed character %q; only letters, marks, numbers, punctuation, symbols, and the encoded space %%20 are allowed"
|
||||
return trace.BadParameter(msg, clip(path), string(r))
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// decode returns the decoded validation view of path s, resolving
|
||||
// only valid escapes. %2F and %2f become "/", %20 a space, and a
|
||||
// non-ASCII escape its byte. The view exposes a structural byte
|
||||
// written as an escape, so "/x%2F..%2Fadmin" is rejected for the same
|
||||
// reason ".." is rejected in "/x/../admin".
|
||||
func decode(s string) string {
|
||||
if !strings.ContainsRune(s, '%') {
|
||||
return s
|
||||
}
|
||||
var b strings.Builder
|
||||
b.Grow(len(s))
|
||||
for i := 0; i < len(s); i++ {
|
||||
if s[i] == '%' && i+2 < len(s) {
|
||||
if v, err := strconv.ParseUint(s[i+1:i+3], 16, 8); err == nil && isAllowedEscape(byte(v)) {
|
||||
b.WriteByte(byte(v))
|
||||
i += 2
|
||||
continue
|
||||
}
|
||||
}
|
||||
b.WriteByte(s[i])
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// rejectDotSegment rejects a segment made of only dots and spaces
|
||||
// that has at least one dot. "." and ".." are traversal segments,
|
||||
// and an upstream that strips spaces or extra dots would resolve
|
||||
// forms like ". ." or "..." the same way.
|
||||
func rejectDotSegment(seg string) error {
|
||||
if strings.Trim(seg, ". ") == "" && strings.Contains(seg, ".") {
|
||||
const msg = `segment %q is only dots and spaces; an upstream could resolve it as "." or ".."`
|
||||
return trace.BadParameter(msg, clip(seg))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// rejectLeadingMark rejects a segment whose first character composes
|
||||
// onto the character before it, so "/a/%CC%87b", where %CC%87 is
|
||||
// U+0307 combining dot above, looks like "/a/b". RFC 5891 bans the
|
||||
// same form at the start of a domain label.
|
||||
func rejectLeadingMark(seg string) error {
|
||||
if seg == "" || seg[0] < utf8.RuneSelf {
|
||||
return nil
|
||||
}
|
||||
r, _ := utf8.DecodeRuneInString(seg)
|
||||
// Some composing characters are not marks, so both checks are needed.
|
||||
if unicode.IsMark(r) || !norm.NFKC.PropertiesString(seg).BoundaryBefore() {
|
||||
const msg = "segment %q starts with %q, which composes onto the character before it; it must follow a base character"
|
||||
return trace.BadParameter(msg, clip(seg), string(r))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// rejectEdgeSpace rejects a segment whose first or last rune is a
|
||||
// space. An upstream that trims the segment would see a different
|
||||
// one, so "..%20" would trim to ".." and "secret%20" to "secret".
|
||||
func rejectEdgeSpace(seg string) error {
|
||||
if strings.HasPrefix(seg, " ") || strings.HasSuffix(seg, " ") {
|
||||
const msg = "segment %q starts or ends with a space; a space must be between other characters"
|
||||
return trace.BadParameter(msg, clip(seg))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// rejectTrailingDot rejects a segment whose trailing run of dots and
|
||||
// spaces contains a dot. IIS and Windows trim trailing dots and
|
||||
// spaces, so an upstream would resolve "secret." and "secret%20." as
|
||||
// "secret", a segment the matcher never saw. A leading dot stays
|
||||
// allowed because paths such as "/.well-known" depend on it.
|
||||
func rejectTrailingDot(seg string) error {
|
||||
trimmed := strings.TrimRight(seg, ". ")
|
||||
if strings.Contains(seg[len(trimmed):], ".") {
|
||||
const msg = "segment %q ends with dots and spaces; an upstream could trim it to %q"
|
||||
return trace.BadParameter(msg, clip(seg), clip(trimmed))
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// isGraphicRune reports whether r is U+0020, a letter, mark, number,
|
||||
// punctuation, or symbol. U+0020 is allowed because it only reaches
|
||||
// the decoded view as %20, and [rejectEdgeSpace] keeps it off the
|
||||
// segment edges. Every other space and separator rune that
|
||||
// unicode.IsGraphic allows is excluded, along with control, format,
|
||||
// surrogate, private-use, and unassigned runes. This is the only
|
||||
// check that rejects the NFKC-stable ones among them, such as U+1680
|
||||
// ogham space mark.
|
||||
func isGraphicRune(r rune) bool {
|
||||
return r == ' ' || unicode.IsLetter(r) || unicode.IsMark(r) ||
|
||||
unicode.IsNumber(r) || unicode.IsPunct(r) || unicode.IsSymbol(r)
|
||||
}
|
||||
|
||||
// clip shortens s for use in an error message. A rejected path can be
|
||||
// kilobytes long, and the message survives into logs and audit events.
|
||||
func clip(s string) string {
|
||||
const limit = 256
|
||||
if len(s) <= limit {
|
||||
return s
|
||||
}
|
||||
return s[:limit] + "..."
|
||||
}
|
||||
@@ -0,0 +1,386 @@
|
||||
/*
|
||||
* Teleport
|
||||
* Copyright (C) 2026 Gravitational, Inc.
|
||||
*
|
||||
* This program is free software: you can redistribute it and/or modify
|
||||
* it under the terms of the GNU Affero General Public License as published by
|
||||
* the Free Software Foundation, either version 3 of the License, or
|
||||
* (at your option) any later version.
|
||||
*
|
||||
* This program is distributed in the hope that it will be useful,
|
||||
* but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
* GNU Affero General Public License for more details.
|
||||
*
|
||||
* You should have received a copy of the GNU Affero General Public License
|
||||
* along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
*/
|
||||
|
||||
package appresource
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"net/url"
|
||||
"strings"
|
||||
"testing"
|
||||
"unicode"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"golang.org/x/text/unicode/norm"
|
||||
"pgregory.net/rapid"
|
||||
)
|
||||
|
||||
// stableMarkPairs are base characters followed by a combining mark
|
||||
// that has no pre-composed form with them, so each pair survives NFKC
|
||||
// and gives the generator a mark where a mark is allowed.
|
||||
var stableMarkPairs = []string{
|
||||
"q̇", // U+0071 U+0307, latin q, combining dot above
|
||||
"v́", // U+0076 U+0301, latin v, combining acute
|
||||
"j̣", // U+006A U+0323, latin j, combining dot below
|
||||
"α̈", // U+03B1 U+0308, greek alpha, combining diaeresis
|
||||
"б́", // U+0431 U+0301, cyrillic be, combining acute
|
||||
"का", // U+0915 U+093E, devanagari ka, vowel sign aa, a spacing mark
|
||||
"ก่", // U+0E01 U+0E48, thai ko kai, mai ek
|
||||
"あ゙", // U+3042 U+3099, hiragana a, combining voiced sound mark
|
||||
"\u05d0\u05b7", // U+05D0 U+05B7, hebrew alef, point patah, escaped because RTL reorders the line
|
||||
"\u0627\u064e", // U+0627 U+064E, arabic alef, fatha, escaped because RTL reorders the line
|
||||
}
|
||||
|
||||
// safePunct is legalPathPunct without the two bytes the generator
|
||||
// places itself, "/" as the segment boundary and "%" as the start of
|
||||
// an escape.
|
||||
var safePunct = strings.NewReplacer("/", "", "%", "").Replace(legalPathPunct)
|
||||
|
||||
const alnum = "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789"
|
||||
|
||||
// TestTokenizePropertyAcceptsWireForm asserts that every path the wire
|
||||
// generator produces is accepted. It pins over-rejection, which the
|
||||
// reject-only properties below cannot see.
|
||||
func TestTokenizePropertyAcceptsWireForm(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
_, err := Tokenize(path)
|
||||
require.NoError(t, err, "generated wire path rejected: %q", path)
|
||||
})
|
||||
}
|
||||
|
||||
// TestTokenizePropertyAcceptsOnlyEscapedForm asserts that an accepted
|
||||
// path is what net/url emits for it. The reverse proxy forwards
|
||||
// EscapedPath, so a path Go would re-encode reaches the app as bytes
|
||||
// the matcher never saw. net/url decides here, not the generator, so
|
||||
// widening a byte rule without widening what net/url preserves fails
|
||||
// this property.
|
||||
func TestTokenizePropertyAcceptsOnlyEscapedForm(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
_, err := Tokenize(path)
|
||||
require.NoError(t, err)
|
||||
u, err := url.ParseRequestURI(path)
|
||||
require.NoError(t, err, "accepted path does not parse: %q", path)
|
||||
require.Equal(t, path, u.EscapedPath(), "accepted path is not its own escaped form")
|
||||
})
|
||||
}
|
||||
|
||||
// TestTokenizePropertyPreservesStructure asserts that tokens are the
|
||||
// input split on raw "/" and nothing else. Rejoining them reproduces
|
||||
// the input byte for byte, so no token was decoded or rewritten.
|
||||
func TestTokenizePropertyPreservesStructure(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
tokens, err := Tokenize(path)
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, path, "/"+strings.Join(tokens, "/"))
|
||||
require.Len(t, tokens, strings.Count(path, "/"))
|
||||
for _, tok := range tokens {
|
||||
require.NotContains(t, tok, "/")
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// TestTokenizePropertyIgnoresHexCase asserts that the hex case of an
|
||||
// escape never changes the verdict, and that the tokens keep the case
|
||||
// they arrived in.
|
||||
func TestTokenizePropertyIgnoresHexCase(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
flipped := flipHexCase(path)
|
||||
tokens, err := Tokenize(path)
|
||||
require.NoError(t, err)
|
||||
flippedTokens, err := Tokenize(flipped)
|
||||
require.NoError(t, err, "hex case flip changed the verdict: %q", flipped)
|
||||
require.Equal(t, flipped, "/"+strings.Join(flippedTokens, "/"))
|
||||
require.Len(t, flippedTokens, len(tokens))
|
||||
})
|
||||
}
|
||||
|
||||
// TestTokenizePropertyRejectsRawNonASCII asserts that a raw byte at or
|
||||
// above 0x80 anywhere in an accepted path is rejected. Unicode must
|
||||
// arrive percent-encoded.
|
||||
func TestTokenizePropertyRejectsRawNonASCII(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
b := byte(rapid.IntRange(0x80, 0xff).Draw(t, "byte"))
|
||||
mutated := insertAt(t, path, string(b))
|
||||
_, err := Tokenize(mutated)
|
||||
require.Error(t, err, "raw byte %#x not rejected: %q", b, mutated)
|
||||
})
|
||||
}
|
||||
|
||||
// TestTokenizePropertyRejectsFoldForms asserts that a known fold or
|
||||
// canonicalization form inserted anywhere in an accepted path is
|
||||
// rejected.
|
||||
func TestTokenizePropertyRejectsFoldForms(t *testing.T) {
|
||||
forms := map[string]string{
|
||||
"%2E (.)": "%2E",
|
||||
"%00 (NUL)": "%00",
|
||||
"%3F (?)": "%3F",
|
||||
"%5C (backslash)": "%5C",
|
||||
"%C0%AF (overlong UTF-8 for /)": "%C0%AF",
|
||||
"%EF%BC%8F (fullwidth solidus)": "%EF%BC%8F",
|
||||
"%E2%80%8B (zero-width space)": "%E2%80%8B",
|
||||
"%C2%A0 (no-break space)": "%C2%A0",
|
||||
"%E1%9A%80 (ogham space mark)": "%E1%9A%80",
|
||||
"%CC%81 (combining acute on café)": "cafe%CC%81",
|
||||
}
|
||||
for name, form := range forms {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
mutated := insertAt(t, path, form)
|
||||
_, err := Tokenize(mutated)
|
||||
require.Error(t, err, "%s not rejected: %q", name, mutated)
|
||||
})
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestTokenizePropertyRejectsBadSegments asserts that a bad segment
|
||||
// spliced in after any "/" of an accepted path is rejected.
|
||||
func TestTokenizePropertyRejectsBadSegments(t *testing.T) {
|
||||
segments := map[string]string{
|
||||
"dot": ".",
|
||||
"dot-dot": "..",
|
||||
"empty": "",
|
||||
"leading mark %CC%87": "%CC%87x",
|
||||
"leading mark after %2F": "x%2F%CC%87y",
|
||||
"encoded dot-dot between": "x%2F..%2Fy",
|
||||
"only a space %20": "%20",
|
||||
"leading space %20": "%20x",
|
||||
"trailing space %20": "x%20",
|
||||
"leading space %20 after %2F": "x%2F%20y",
|
||||
"trailing space %20 before %2F": "x%20%2Fy",
|
||||
"dot space dot": ".%20.",
|
||||
"dot-dot space dot": "..%20.",
|
||||
"only dots": "...",
|
||||
"dots and spaces after %2F": "x%2F.%20.",
|
||||
"trailing dot": "x.",
|
||||
"trailing dots": "x..",
|
||||
"trailing space dot": "x%20.",
|
||||
"trailing dot before %2F": "x.%2Fy",
|
||||
}
|
||||
for name, segment := range segments {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
var slashes []int
|
||||
for i := range len(path) {
|
||||
if path[i] == '/' {
|
||||
slashes = append(slashes, i)
|
||||
}
|
||||
}
|
||||
n := rapid.IntRange(0, len(slashes)-1).Draw(t, "slash")
|
||||
i := slashes[n]
|
||||
mutated := path[:i+1] + segment + "/" + path[i+1:]
|
||||
_, err := Tokenize(mutated)
|
||||
require.Error(t, err, "segment %q not rejected: %q", segment, mutated)
|
||||
})
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestTokenizePropertyRejectsEdgeSpaces asserts that a %20 spliced at
|
||||
// the start or end of any part of an accepted path is rejected. An
|
||||
// upstream that trims the part would see a different part.
|
||||
func TestTokenizePropertyRejectsEdgeSpaces(t *testing.T) {
|
||||
rapid.Check(t, func(t *rapid.T) {
|
||||
path := wirePathGen().Draw(t, "path")
|
||||
starts, ends := partBounds(path)
|
||||
bounds := starts
|
||||
if rapid.Bool().Draw(t, "trailing") {
|
||||
bounds = ends
|
||||
}
|
||||
i := rapid.SampledFrom(bounds).Draw(t, "bound")
|
||||
mutated := path[:i] + "%20" + path[i:]
|
||||
_, err := Tokenize(mutated)
|
||||
require.Error(t, err, "edge space not rejected: %q", mutated)
|
||||
})
|
||||
}
|
||||
|
||||
// wirePathGen produces a valid wire path, one or more segments with an
|
||||
// optional trailing slash.
|
||||
func wirePathGen() *rapid.Generator[string] {
|
||||
return rapid.Custom(func(t *rapid.T) string {
|
||||
n := rapid.IntRange(1, 4).Draw(t, "segment-count")
|
||||
segments := make([]string, n)
|
||||
for i := range segments {
|
||||
segments[i] = segmentGen().Draw(t, "segment")
|
||||
}
|
||||
path := "/" + strings.Join(segments, "/")
|
||||
if rapid.Bool().Draw(t, "trailing-slash") {
|
||||
path += "/"
|
||||
}
|
||||
return path
|
||||
})
|
||||
}
|
||||
|
||||
// segmentGen joins one or more parts with the encoded separator, %2F
|
||||
// or %2f.
|
||||
func segmentGen() *rapid.Generator[string] {
|
||||
return rapid.Custom(func(t *rapid.T) string {
|
||||
n := rapid.IntRange(1, 3).Draw(t, "part-count")
|
||||
parts := make([]string, n)
|
||||
for i := range parts {
|
||||
parts[i] = partGen().Draw(t, "part")
|
||||
}
|
||||
sep := rapid.SampledFrom([]string{"%2F", "%2f"}).Draw(t, "encoded-separator")
|
||||
return strings.Join(parts, sep)
|
||||
})
|
||||
}
|
||||
|
||||
// partGen produces the percent-encoded content between two separators.
|
||||
// No part starts with a mark and no rune composes with its neighbor.
|
||||
// A space, emitted as %20, composes with no neighbor under NFKC, so
|
||||
// it needs no neighbor rule. The filter drops parts made of only
|
||||
// dots and spaces, parts that start or end with a space, and parts
|
||||
// whose trailing run of dots and spaces contains a dot.
|
||||
func partGen() *rapid.Generator[string] {
|
||||
content := rapid.Custom(func(t *rapid.T) string {
|
||||
var b strings.Builder
|
||||
n := rapid.IntRange(1, 5).Draw(t, "rune-count")
|
||||
for range n {
|
||||
switch rapid.IntRange(0, 4).Draw(t, "rune-kind") {
|
||||
case 0:
|
||||
b.WriteByte(alnum[rapid.IntRange(0, len(alnum)-1).Draw(t, "alnum")])
|
||||
case 1:
|
||||
b.WriteByte(safePunct[rapid.IntRange(0, len(safePunct)-1).Draw(t, "punct")])
|
||||
case 2:
|
||||
b.WriteRune(safeRuneGen().Draw(t, "non-ascii"))
|
||||
case 3:
|
||||
b.WriteString(rapid.SampledFrom(stableMarkPairs).Draw(t, "mark-pair"))
|
||||
case 4:
|
||||
b.WriteByte(' ')
|
||||
}
|
||||
if rapid.Bool().Draw(t, "trailing-mark") {
|
||||
b.WriteRune(freeMarkGen().Draw(t, "mark"))
|
||||
}
|
||||
}
|
||||
return b.String()
|
||||
}).Filter(func(s string) bool {
|
||||
return strings.Trim(s, ". ") != "" &&
|
||||
!strings.HasPrefix(s, " ") && !strings.HasSuffix(s, " ") &&
|
||||
!strings.Contains(s[len(strings.TrimRight(s, ". ")):], ".")
|
||||
})
|
||||
return rapid.Custom(func(t *rapid.T) string {
|
||||
return encodeNonASCII(t, content.Draw(t, "content"))
|
||||
})
|
||||
}
|
||||
|
||||
// safeRuneGen produces a non-ASCII rune that is legal anywhere in a
|
||||
// part, including first, so it excludes the marks.
|
||||
func safeRuneGen() *rapid.Generator[rune] {
|
||||
return freeRuneGen().Filter(func(r rune) bool { return !unicode.IsMark(r) })
|
||||
}
|
||||
|
||||
// freeMarkGen produces a mark that stands on its own, so it is legal
|
||||
// anywhere except as the first rune of a part.
|
||||
func freeMarkGen() *rapid.Generator[rune] {
|
||||
return freeRuneGen().Filter(unicode.IsMark)
|
||||
}
|
||||
|
||||
// freeRuneGen produces a non-ASCII rune that NFKC puts a boundary
|
||||
// before, so it never combines with whatever the generator wrote
|
||||
// previously. That covers the combining marks with a non-zero class
|
||||
// and the trailing Hangul jamo without naming either. The filter reads
|
||||
// the stdlib category tables and the normalization properties rather
|
||||
// than isGraphicRune, so the generator does not inherit the
|
||||
// validator's own view of what is allowed.
|
||||
func freeRuneGen() *rapid.Generator[rune] {
|
||||
return rapid.Rune().Filter(func(r rune) bool {
|
||||
if r < 0x80 || !unicode.IsGraphic(r) || unicode.IsSpace(r) {
|
||||
return false
|
||||
}
|
||||
return norm.NFKC.PropertiesString(string(r)).BoundaryBefore() &&
|
||||
norm.NFKC.IsNormalString(string(r))
|
||||
})
|
||||
}
|
||||
|
||||
// encodeNonASCII percent-encodes the bytes of s at or above 0x80 and
|
||||
// the space, which has no raw form in a path, and leaves the other
|
||||
// ASCII bytes raw, in the hex case the generator picks.
|
||||
func encodeNonASCII(t *rapid.T, s string) string {
|
||||
format := "%%%02x"
|
||||
if rapid.Bool().Draw(t, "upper-hex") {
|
||||
format = "%%%02X"
|
||||
}
|
||||
var b strings.Builder
|
||||
for i := range len(s) {
|
||||
if s[i] < 0x80 && s[i] != ' ' {
|
||||
b.WriteByte(s[i])
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(&b, format, s[i])
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// partBounds returns the offsets where a part starts and ends. A part
|
||||
// is bounded by a raw slash, an encoded separator, or the end of the
|
||||
// path.
|
||||
func partBounds(path string) (starts, ends []int) {
|
||||
for i := range len(path) {
|
||||
switch {
|
||||
case path[i] == '/':
|
||||
starts = append(starts, i+1)
|
||||
if i > 0 {
|
||||
ends = append(ends, i)
|
||||
}
|
||||
case path[i] == '%' && i+2 < len(path) && path[i+1] == '2' && (path[i+2] == 'F' || path[i+2] == 'f'):
|
||||
starts = append(starts, i+3)
|
||||
ends = append(ends, i)
|
||||
}
|
||||
}
|
||||
ends = append(ends, len(path))
|
||||
return starts, ends
|
||||
}
|
||||
|
||||
// insertAt splices s into path at a generated offset after the leading
|
||||
// slash.
|
||||
func insertAt(t *rapid.T, path, s string) string {
|
||||
i := rapid.IntRange(1, len(path)).Draw(t, "offset")
|
||||
return path[:i] + s + path[i:]
|
||||
}
|
||||
|
||||
// flipHexCase swaps the case of the hex digits of every escape in
|
||||
// path.
|
||||
func flipHexCase(path string) string {
|
||||
out := []byte(path)
|
||||
for i := 0; i+2 < len(out); i++ {
|
||||
if out[i] != '%' {
|
||||
continue
|
||||
}
|
||||
out[i+1] = flipByteCase(out[i+1])
|
||||
out[i+2] = flipByteCase(out[i+2])
|
||||
}
|
||||
return string(out)
|
||||
}
|
||||
|
||||
// flipByteCase swaps the case of an ASCII letter.
|
||||
func flipByteCase(b byte) byte {
|
||||
switch {
|
||||
case b >= 'a' && b <= 'z':
|
||||
return b - 'a' + 'A'
|
||||
case b >= 'A' && b <= 'Z':
|
||||
return b - 'A' + 'a'
|
||||
}
|
||||
return b
|
||||
}
|
||||
@@ -0,0 +1,568 @@
|
||||
/*
|
||||
* Teleport
|
||||
* Copyright (C) 2026 Gravitational, Inc.
|
||||
*
|
||||
* This program is free software: you can redistribute it and/or modify
|
||||
* it under the terms of the GNU Affero General Public License as published by
|
||||
* the Free Software Foundation, either version 3 of the License, or
|
||||
* (at your option) any later version.
|
||||
*
|
||||
* This program is distributed in the hope that it will be useful,
|
||||
* but WITHOUT ANY WARRANTY; without even the implied warranty of
|
||||
* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
|
||||
* GNU Affero General Public License for more details.
|
||||
*
|
||||
* You should have received a copy of the GNU Affero General Public License
|
||||
* along with this program. If not, see <http://www.gnu.org/licenses/>.
|
||||
*/
|
||||
|
||||
package appresource
|
||||
|
||||
import (
|
||||
"net/url"
|
||||
"strings"
|
||||
"testing"
|
||||
"unicode"
|
||||
"unicode/utf8"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"golang.org/x/text/unicode/norm"
|
||||
)
|
||||
|
||||
// TestTokenize pins the tokenizer's accept and reject cases, including
|
||||
// the opaque encoded separator and the decode-for-validation view.
|
||||
func TestTokenize(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
path string
|
||||
want []string
|
||||
wantErr bool
|
||||
}{
|
||||
{
|
||||
name: "plain path splits on real slashes",
|
||||
path: "/api/v4/projects",
|
||||
want: []string{"api", "v4", "projects"},
|
||||
},
|
||||
{
|
||||
name: "bare root yields a single empty token",
|
||||
path: "/",
|
||||
want: []string{""},
|
||||
},
|
||||
{
|
||||
name: "path at the length cap is allowed",
|
||||
path: "/" + strings.Repeat("a", maxPathLength-1),
|
||||
want: []string{strings.Repeat("a", maxPathLength-1)},
|
||||
},
|
||||
{
|
||||
name: "encoded slash stays one opaque token",
|
||||
path: "/files/a%2Fb",
|
||||
want: []string{"files", "a%2Fb"},
|
||||
},
|
||||
{
|
||||
name: "lowercase encoded slash stays one raw token",
|
||||
path: "/files/a%2fb",
|
||||
want: []string{"files", "a%2fb"},
|
||||
},
|
||||
{
|
||||
name: "trailing slash yields a trailing empty token",
|
||||
path: "/files/",
|
||||
want: []string{"files", ""},
|
||||
},
|
||||
{
|
||||
name: "trailing encoded slash is allowed raw",
|
||||
path: "/files/a%2F",
|
||||
want: []string{"files", "a%2F"},
|
||||
},
|
||||
{
|
||||
// é arrives percent-encoded and stays raw in the token.
|
||||
name: "percent-encoded UTF-8 content is allowed raw",
|
||||
path: "/files/caf%C3%A9.md",
|
||||
want: []string{"files", "caf%C3%A9.md"},
|
||||
},
|
||||
{
|
||||
name: "encoded space %20 stays raw in the token",
|
||||
path: "/job/My%20Job/lastBuild",
|
||||
want: []string{"job", "My%20Job", "lastBuild"},
|
||||
},
|
||||
{
|
||||
name: "encoded space %20 in the first segment",
|
||||
path: "/My%20Jobs/config",
|
||||
want: []string{"My%20Jobs", "config"},
|
||||
},
|
||||
{
|
||||
name: "encoded space %20 in the last segment",
|
||||
path: "/job/last%20build",
|
||||
want: []string{"job", "last%20build"},
|
||||
},
|
||||
{
|
||||
name: "several encoded spaces %20 in one segment",
|
||||
path: "/job/a%20b%20c",
|
||||
want: []string{"job", "a%20b%20c"},
|
||||
},
|
||||
{
|
||||
name: "consecutive encoded spaces %20%20 are allowed",
|
||||
path: "/job/a%20%20b",
|
||||
want: []string{"job", "a%20%20b"},
|
||||
},
|
||||
{
|
||||
name: "encoded spaces %20 in several segments are allowed",
|
||||
path: "/sites/Team%20Site/Shared%20Documents/x",
|
||||
want: []string{"sites", "Team%20Site", "Shared%20Documents", "x"},
|
||||
},
|
||||
{
|
||||
name: "interior encoded space %20 next to the encoded separator %2F",
|
||||
path: "/files/a%20b%2Fc",
|
||||
want: []string{"files", "a%20b%2Fc"},
|
||||
},
|
||||
{
|
||||
name: "encoded space %20 next to a non-ASCII escape",
|
||||
path: "/files/caf%C3%A9%20menu",
|
||||
want: []string{"files", "caf%C3%A9%20menu"},
|
||||
},
|
||||
{
|
||||
name: "dots and an encoded space %20 among other characters are allowed",
|
||||
path: "/a/a.%20.b/c",
|
||||
want: []string{"a", "a.%20.b", "c"},
|
||||
},
|
||||
{
|
||||
name: "combining mark %CC%87 (U+0307) on a base character is allowed",
|
||||
path: "/p/q%CC%87x",
|
||||
want: []string{"p", "q%CC%87x"},
|
||||
},
|
||||
{
|
||||
name: "interior dots in a segment are allowed",
|
||||
path: "/a/1.2.3/b",
|
||||
want: []string{"a", "1.2.3", "b"},
|
||||
},
|
||||
{
|
||||
name: "a leading dot in a segment is allowed",
|
||||
path: "/.well-known/openid-configuration",
|
||||
want: []string{".well-known", "openid-configuration"},
|
||||
},
|
||||
{
|
||||
// The segment ends with the mark riding the space, not
|
||||
// with a space byte, so no upstream trim fires on it.
|
||||
name: "a combining mark %CC%81 on a trailing interior space %20 is allowed",
|
||||
path: "/p/x%20%CC%81",
|
||||
want: []string{"p", "x%20%CC%81"},
|
||||
},
|
||||
{
|
||||
name: "path over the length cap is rejected",
|
||||
path: "/" + strings.Repeat("a", maxPathLength),
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "double-encoded slash is rejected because %25 decodes to %",
|
||||
path: "/files/a%252Fb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an ASCII escape %40 (@) is rejected",
|
||||
path: "/files/a%40b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded dot %2E (.) is rejected",
|
||||
path: "/files/a%2Eb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded hash %23 (#) is rejected",
|
||||
path: "/files/a%23b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded percent %25 (%) is rejected",
|
||||
path: "/files/a%25b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded question mark %3F (?) is rejected",
|
||||
path: "/files/a%3Fb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded NUL %00 is rejected",
|
||||
path: "/files/a%00b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded double quote %22 (\") is rejected",
|
||||
path: "/files/a%22b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded less-than %3C (<) is rejected",
|
||||
path: "/files/a%3Cb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded greater-than %3E (>) is rejected",
|
||||
path: "/files/a%3Eb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded opening bracket %5B ([) is rejected",
|
||||
path: "/files/a%5Bb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded closing bracket %5D (]) is rejected",
|
||||
path: "/files/a%5Db",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded backslash %5C (\\) is rejected",
|
||||
path: "/files/a%5Cb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded caret %5E (^) is rejected",
|
||||
path: "/files/a%5Eb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded backtick %60 (`) is rejected",
|
||||
path: "/files/a%60b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded opening brace %7B ({) is rejected",
|
||||
path: "/files/a%7Bb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded pipe %7C (|) is rejected",
|
||||
path: "/files/a%7Cb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an encoded closing brace %7D (}) is rejected",
|
||||
path: "/files/a%7Db",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a segment of only the encoded space %20 is rejected",
|
||||
path: "/a/%20/b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a leading encoded space %20 in a segment is rejected",
|
||||
path: "/a/%20x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a trailing encoded space %20 in a segment is rejected",
|
||||
path: "/a/x%20",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
// The segment decodes to ".. ", which trims to "..".
|
||||
name: "a dot-dot with a trailing encoded space %20 is rejected",
|
||||
path: "/a/..%20/b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
// The segment decodes to ". ", which trims to ".".
|
||||
name: "a dot with a trailing encoded space %20 is rejected",
|
||||
path: "/a/.%20/b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a dot-dot with a leading encoded space %20 is rejected",
|
||||
path: "/a/%20../b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a leading encoded space %20 after an encoded slash %2F is rejected",
|
||||
path: "/p/a%2F%20b/c",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a trailing encoded space %20 before an encoded slash %2F is rejected",
|
||||
path: "/p/b%20%2Fc/d",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
// The segment starts with a space that carries the mark.
|
||||
name: "a combining mark %CC%81 (U+0301) on a leading encoded space %20 is rejected",
|
||||
path: "/p/%20%CC%81x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
// The segment decodes to ". .", stripped of spaces "..".
|
||||
name: "an encoded space %20 between dots is rejected",
|
||||
path: "/a/.%20./b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
// The dot-segment rule fires; the segment also violates
|
||||
// the edge-space rule.
|
||||
name: "a segment of alternating %20 and dots is rejected",
|
||||
path: "/a/%20.%20.%20/b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a dot-dot with %20 and a trailing dot is rejected",
|
||||
path: "/a/..%20./b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "dots separated by encoded spaces %20 are rejected",
|
||||
path: "/a/.%20.%20./b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "dots and %20 between encoded slashes %2F are rejected",
|
||||
path: "/p/a%2F.%20.%2Fb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a segment of only dots is rejected",
|
||||
path: "/a/.../b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a trailing dot in a segment is rejected",
|
||||
path: "/files/secret./x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "trailing dots in a segment are rejected",
|
||||
path: "/files/secret../x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a trailing encoded space %20 and dot in a segment are rejected",
|
||||
path: "/files/secret%20./x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a trailing dot before an encoded slash %2F is rejected",
|
||||
path: "/p/a.%2Fb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a truncated escape is rejected",
|
||||
path: "/files/a%2",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a malformed escape with non-hex digits is rejected",
|
||||
path: "/files/a%G1b",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a lone percent is rejected",
|
||||
path: "/files/a%",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a dot-dot between encoded slashes is rejected",
|
||||
path: "/a%2F..%2Fadmin",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "an empty inner part in encoded slashes is rejected",
|
||||
path: "/a%2F%2Fb",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a raw dot-dot segment is rejected",
|
||||
path: "/api/v4/../secret",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a raw single-dot segment is rejected",
|
||||
path: "/api/./v4",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a segment starting with the conjoining jamo %E1%86%A8 (U+11A8) is rejected",
|
||||
path: "/%EA%B0%80/%E1%86%A8",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a segment starting with the combining mark %CC%87 is rejected",
|
||||
path: "/p/%CC%87x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a combining mark after an encoded slash %2F is rejected",
|
||||
path: "/p/a%2F%CC%87x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "consecutive slashes are rejected",
|
||||
path: "/api//v4",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "a path without a leading slash is rejected",
|
||||
path: "api/v4",
|
||||
wantErr: true,
|
||||
},
|
||||
}
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
got, err := Tokenize(tt.path)
|
||||
if tt.wantErr {
|
||||
require.Error(t, err)
|
||||
return
|
||||
}
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, tt.want, got)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestTokenizeByteRules pins the RFC 3986 path-character rules.
|
||||
// Pchar except for ";", plus "/" and "%", are allowed. Anything else
|
||||
// in the raw path is rejected.
|
||||
func TestTokenizeByteRules(t *testing.T) {
|
||||
allow := []string{
|
||||
"/api/v4/projects",
|
||||
"/api/@@@", // "@" is a pchar
|
||||
"/api/(group)/sub.tree", // sub-delims and unreserved
|
||||
"/api/a:b/c,d/e=f/gh", // ":" and sub-delims
|
||||
"/api/a-b_c~d!$&'()*+,=", // the unreserved and sub-delim set, except for ";"
|
||||
"/api/v4/projects/group%2Frepo/x", // the encoded separator, allowed
|
||||
}
|
||||
for _, p := range allow {
|
||||
t.Run("allow "+p, func(t *testing.T) {
|
||||
_, err := Tokenize(p)
|
||||
require.NoError(t, err)
|
||||
})
|
||||
}
|
||||
|
||||
reject := []string{
|
||||
"/api/a b", // raw space, allowed only encoded as %20
|
||||
"/api/a\"b", // double quote
|
||||
"/api/a<b", // angle bracket
|
||||
"/api/a>b", // angle bracket
|
||||
"/api/a{b}", // braces
|
||||
"/api/a|b", // pipe
|
||||
"/api/a^b", // caret
|
||||
"/api/a`b", // backtick
|
||||
"/api/a\\b", // backslash
|
||||
"/api/a[b]", // square brackets
|
||||
"/api/a#b", // fragment delimiter
|
||||
"/api/a?b", // query delimiter
|
||||
"/api/café", // raw non-ASCII, must be percent-encoded
|
||||
"/api/a;b/c", // semicolon, the matrix-parameter / jsessionid vector
|
||||
}
|
||||
for _, p := range reject {
|
||||
t.Run("reject "+p, func(t *testing.T) {
|
||||
_, err := Tokenize(p)
|
||||
require.Error(t, err)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestNonASCIIFold pins the fold and homoglyph cases for the
|
||||
// non-ASCII path pipeline. Rejected entries are forms that could
|
||||
// resolve to a different segment once an upstream normalizes them.
|
||||
func TestNonASCIIFold(t *testing.T) {
|
||||
allow := map[string]string{
|
||||
"precomposed accent (café.md)": "/files/caf%C3%A9.md",
|
||||
"CJK han character": "/files/%E6%97%A5.txt",
|
||||
"cyrillic letter": "/u/%D0%B4",
|
||||
"emoji is a symbol": "/r/%F0%9F%98%80",
|
||||
"accent next to encoded slash": "/p/caf%C3%A9%2Fx",
|
||||
}
|
||||
for name, path := range allow {
|
||||
t.Run("allow "+name, func(t *testing.T) {
|
||||
_, err := Tokenize(path)
|
||||
require.NoError(t, err)
|
||||
})
|
||||
}
|
||||
|
||||
reject := map[string]struct {
|
||||
path string
|
||||
// errContains pins which rule must reject inputs that more
|
||||
// than one layer would catch. When empty, any error passes.
|
||||
errContains string
|
||||
}{
|
||||
"raw non-ASCII byte": {path: "/files/caf\xc3\xa9"},
|
||||
"overlong UTF-8 of slash": {path: "/p/a%C0%AFb"},
|
||||
"lone continuation byte": {path: "/p/%A9"},
|
||||
"truncated two-byte sequence": {path: "/p/%C3"},
|
||||
"fullwidth solidus folds to /": {path: "/p/a%EF%BC%8Fb"},
|
||||
"fullwidth A folds to A": {path: "/p/%EF%BC%A1dmin"},
|
||||
"fullwidth lowercase a": {path: "/p/%EF%BD%81dmin"},
|
||||
"zero-width space is format": {path: "/p/a%E2%80%8Bb"},
|
||||
"bidi override is format": {path: "/p/a%E2%80%AEb"},
|
||||
"decomposed e plus accent": {path: "/p/cafe%CC%81"},
|
||||
"ligature fi folds to fi": {path: "/p/o%EF%AC%81ce"},
|
||||
"non-breaking space folds": {path: "/p/a%C2%A0b", errContains: "not NFKC-normalized"},
|
||||
"en quad %E2%80%80 (U+2000) folds to space": {path: "/p/a%E2%80%80b", errContains: "not NFKC-normalized"},
|
||||
"ideographic space %E3%80%80 (U+3000) folds": {path: "/p/a%E3%80%80b", errContains: "not NFKC-normalized"},
|
||||
"ogham space mark %E1%9A%80 (U+1680) is a separator": {path: "/p/a%E1%9A%80b", errContains: "disallowed character"},
|
||||
"line separator %E2%80%A8 (U+2028) is a separator": {path: "/p/a%E2%80%A8b", errContains: "disallowed character"},
|
||||
"conjoining jamo %E1%86%A8 (U+11A8) composes": {path: "/p/%EA%B0%80%E1%86%A8", errContains: "not NFKC-normalized"},
|
||||
"conjoining jamo starting a segment": {path: "/%EA%B0%80/%E1%86%A8", errContains: "composes onto the character before it"},
|
||||
}
|
||||
for name, tt := range reject {
|
||||
t.Run("reject "+name, func(t *testing.T) {
|
||||
_, err := Tokenize(tt.path)
|
||||
require.Error(t, err)
|
||||
if tt.errContains != "" {
|
||||
require.ErrorContains(t, err, tt.errContains)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// FuzzTokenizeNonASCII checks that every accepted path decodes per
|
||||
// segment to valid, NFKC-stable UTF-8. Corpus seeds cover the known
|
||||
// fold and homoglyph bypasses. The fuzzer explores around them.
|
||||
func FuzzTokenizeNonASCII(f *testing.F) {
|
||||
for _, seed := range []string{
|
||||
"/files/caf%C3%A9.md", "/p/a%EF%BC%8Fb", "/p/%EF%BC%A1dmin",
|
||||
"/p/a%C0%AFb", "/p/a%E2%80%8Bb", "/p/cafe%CC%81", "/api/v4/x%2Fy",
|
||||
"/job/My%20Job/lastBuild", "/p/%20%CC%81x", "/p/a%C2%A0b",
|
||||
"/a/..%20/b", "/a/%20/b", "/a/.%20./b", "/a/.../b",
|
||||
"/%EA%B0%80/%E1%86%A8",
|
||||
"/files/secret./x", "/files/secret%20./x",
|
||||
} {
|
||||
f.Add(seed)
|
||||
}
|
||||
f.Fuzz(func(t *testing.T, path string) {
|
||||
tokens, err := Tokenize(path)
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
for _, tok := range tokens {
|
||||
content := decode(tok)
|
||||
valid := utf8.ValidString(content)
|
||||
require.True(t, valid, "accepted token %q decodes to invalid UTF-8", tok)
|
||||
normal := norm.NFKC.IsNormalString(content)
|
||||
require.True(t, normal, "accepted token %q is not NFKC-stable; a fold bypass slipped through", tok)
|
||||
for _, part := range strings.Split(content, "/") {
|
||||
if part == "" {
|
||||
continue
|
||||
}
|
||||
r, _ := utf8.DecodeRuneInString(part)
|
||||
require.False(t, unicode.IsMark(r), "accepted token %q has a part starting with the combining mark %q", tok, string(r))
|
||||
require.False(t, strings.HasPrefix(part, " "), "accepted token %q has a part starting with a space", tok)
|
||||
require.False(t, strings.HasSuffix(part, " "), "accepted token %q has a part ending with a space", tok)
|
||||
dotSpaces := strings.Trim(part, ". ") == "" && strings.Contains(part, ".")
|
||||
require.False(t, dotSpaces, "accepted token %q has a part of only dots and spaces", tok)
|
||||
tail := part[len(strings.TrimRight(part, ". ")):]
|
||||
require.NotContains(t, tail, ".", "accepted token %q has a part ending with dots and spaces", tok)
|
||||
}
|
||||
}
|
||||
// An accepted token contains no raw non-ASCII bytes.
|
||||
for _, tok := range tokens {
|
||||
for i := range len(tok) {
|
||||
require.Less(t, tok[i], byte(0x80), "accepted token %q has a raw non-ASCII byte", tok)
|
||||
}
|
||||
}
|
||||
// Rejoining path segments roundtrips cleanly.
|
||||
require.Equal(t, path, "/"+strings.Join(tokens, "/"))
|
||||
// An accepted path is its own escaped form.
|
||||
u, err := url.ParseRequestURI(path)
|
||||
require.NoError(t, err, "accepted path does not parse: %q", path)
|
||||
require.Equal(t, path, u.EscapedPath(), "accepted path is not its own escaped form")
|
||||
})
|
||||
}
|
||||
Reference in New Issue
Block a user