Validate ULID checkpoint IDs via oklog/ulid · Entire

Validate ULID checkpoint IDs via oklog/ulid

a75a7a8→main·

Soph·2w ago·4 files·+28 added/-7 removed

Replace the hand-rolled Crockford-base32 regex check in KindOf with oklog/ulid's ParseStrict, the same library we'll use to generate ULIDs — so validation and generation agree by construction and we stop maintaining the alphabet by hand. It is also stricter than the char class: it rejects a timestamp overflow (first character > 7).

Canonical form only: isULID also requires v.String() == s, so lowercase and Crockford-normalized aliases (I/L→1, O→0) that ParseStrict tolerates are rejected — we only ever emit the 26-char uppercase canonical encoding, and keeping stored IDs canonical keeps ref names/shards deterministic.

ULIDPattern stays as the loose extraction shape for scanning the Entire-Checkpoint trailer (a regex is needed to find a candidate token in free text); the captured token is then authoritatively validated through this path. The now-unused ulidRegex is removed. Adds the overflow case to the KindOf table.

Sessions

175a19bd932bView transcript

[?
Build Checkpoints Store Based on DesignClaude Code·1 step](/content/gh/entireio/cli/session/6852b33a-0d22-4364-aa6c-8de706ecc215#timeline-175a19bd932b/index.html)

Changes

4

7 unmodified lines

8
9
10
11
12
13
14
15
13 unmodified lines

29
30
31
30
31
32
33
32
33
34
35
36
37
38
39
40
41
8 unmodified lines

50
51
52
48
49
53
54
55
56
57
58
59
60
61
62
63
64
65
66
14 unmodified lines

81
82
83
70
84
85
86
87

7 unmodified lines

"encoding/json"
    "fmt"
    "regexp"

ulid "github.com/oklog/ulid/v2"

// CheckpointID identifies a checkpoint. It comes in two formats: a legacy
13 unmodified lines

// CheckpointPattern for matching a checkpoint ID that may be either format.
const Pattern = `[0-9a-f]{12}`

// ULIDPattern is the regex pattern for a ULID checkpoint ID: exactly 26 Crockford
// base32 characters (digits plus uppercase A-Z excluding I, L, O, U). Future
// checkpoint IDs are expected to be ULIDs so they sort lexicographically by
// creation time.
// ULIDPattern is the regex SHAPE of a ULID checkpoint ID: 26 Crockford base32
// characters (digits plus uppercase A-Z excluding I, L, O, U). It exists to
// extract a candidate checkpoint ID from free text (see CheckpointPattern); it
// is intentionally looser than the authoritative validation in KindOf/Validate,
// which decode the value via oklog/ulid (rejecting e.g. a timestamp overflow the
// char class alone would accept). Future checkpoint IDs are expected to be ULIDs
// so they sort lexicographically by creation time.
const ULIDPattern = `[0-9ABCDEFGHJKMNPQRSTVWXYZ]{26}`

// CheckpointPattern matches a checkpoint ID in free text in either format
8 unmodified lines

// checkpointIDRegex validates the legacy format: exactly 12 lowercase hex characters.
var checkpointIDRegex = regexp.MustCompile(`^` + Pattern + `$`)

// ulidRegex validates the ULID format: exactly 26 Crockford base32 characters.
var ulidRegex = regexp.MustCompile(`^` + ULIDPattern + `$`)
// isULID reports whether s is a ULID in canonical form, decoded via oklog/ulid —
// the same library that will generate ULIDs — so validation and generation agree
// by construction. ParseStrict enforces the 26-char length, the Crockford
// alphabet, and the timestamp-overflow bound (first character must be 0-7). The
// round-trip (v.String() == s) additionally requires the canonical uppercase
// encoding: it rejects lowercase and Crockford-normalized aliases (e.g. I/L→1,
// O→0) that ParseStrict would otherwise accept but we never emit.
func isULID(s string) bool {
    v, err := ulid.ParseStrict(s)
    return err == nil && v.String() == s
}

// Kind classifies a checkpoint ID by its storage format. The two valid kinds
// shard differently when stored as a git ref (see ShardFor).
14 unmodified lines

switch {
    case checkpointIDRegex.MatchString(s):
        return KindLegacy
    case ulidRegex.MatchString(s):
    case isULID(s):
        return KindULID
    default:
        return KindUnknown
}

Mcmd/entire/cli/checkpoint/id/id.go+21/-7

125 unmodified lines

126
127
128
129
130
131
132
133
134

125 unmodified lines

{"legacy all digits", "012345678901", KindLegacy},
        {"ulid", sampleULID, KindULID},
        {"ulid all valid base32", "0123456789ABCDEFGHJKMNPQRS", KindULID},
        // Right charset/length but the timestamp overflows (first char > 7);
        // oklog/ulid rejects it where a plain char-class regex would not.
        {"ulid timestamp overflow", "8123456789ABCDEFGHJKMNPQRS", KindUnknown},
        {"uppercase hex is not legacy", "A1B2C3D4E5F6", KindUnknown},
        {"ulid wrong length", "01KVBJCWYA4YW6J5M9GP655HZ", KindUnknown},
        {"ulid with excluded I", "01KVBJCWYA4YW6J5M9GP655HZI", KindUnknown},

Mcmd/entire/cli/checkpoint/id/id_test.go+3

24 unmodified lines

25
26
27
28
29
30
31

24 unmodified lines

github.com/mattn/go-runewidth v0.0.24
    github.com/muesli/termenv v0.16.0
    github.com/ogen-go/ogen v1.22.0
    github.com/oklog/ulid/v2 v2.1.1
    github.com/posthog/posthog-go v1.16.2
    github.com/sergi/go-diff v1.4.0
    github.com/spf13/cobra v1.10.2

Mgo.mod+1

229 unmodified lines

230
231
232
233
234
235
236
237
238

229 unmodified lines

github.com/nwaples/rardecode/v2 v2.2.2/go.mod h1:7uz379lSxPe6j9nvzxUZ+n7mnJNgjsRNb6IbvGVHRmw=
github.com/ogen-go/ogen v1.22.0 h1:7wU+jcIKg/JBAhM95909ULLdAkGr43KQOuvNpJ7Mxb4=
github.com/ogen-go/ogen v1.22.0/go.mod h1:7BOh9a51QiPCC92RMrj1LlkLjejhBAyPhR+oMc6lR9g=
github.com/oklog/ulid/v2 v2.1.1 h1:suPZ4ARWLOJLegGFiZZ1dFAkqzhMjL3J1TzI+5wHz8s=
github.com/oklog/ulid/v2 v2.1.1/go.mod h1:rcEKHmBBKfef9DhnvX7y1HZBYxjXb0cP5ExxNsTT1QQ=
github.com/pborman/getopt v0.0.0-20170112200414-7148bc3a4c30/go.mod h1:85jBQOZwpVEaDAr341tbn15RS4fCAsIst0qp7i8ex1o=
github.com/pelletier/go-toml/v2 v2.3.1 h1:MYEvvGnQjeNkRF1qUuGolNtNExTDwct51yp7olPtrEc=
github.com/pelletier/go-toml/v2 v2.3.1/go.mod h1:2gIqNv+qfxSVS7cM2xJQKtLSTLUE9V8t9Stt+h56mCY=
github.com/pierrec/lz4/v4 v4.1.26 h1:GrpZw1gZttORinvzBdXPUXATeqlJjqUG/D87TKMnhjY=

Mgo.sum+3