Austin, drawn from real OpenStreetMap data, on a hex lattice where every cell is a self-describing GYST UUIDv8.
Live: austinsays.com/city · austinsays.com/about
A craigslist/nextdoor-shaped city site, built on one idea: a place should have an address you can decode.
- 29,644 real OSM building footprints, triangulated, drawn in one WebGL call
- Each building's centroid → a pointy-top hex cell → a GYST UUIDv8 (type
0x340) that decodes back to(q, r)with no database and no lookup table - A feed of city chatter run through Jev (TypeSafe's System One model), where each post either earns a durable identity or explicitly does not
The map is not a picture of Austin. It is an index of Austin.
lat/lon → (q, r) → UUIDv8 → back to (q, r)
// lib/hexgrid-austin.ts
AUSTIN_FRAME = { id:'austin', origin:{lat:30.2672, lon:-97.7431}, cellKm:2, ns:0x9d1 }
// ns 0x9d1 = FNV-1a-12 of "austin" — so cells cannot alias across citiesPointy-top axial hexes, x = size·√3·(q + r/2), y = size·1.5·r.
The hex is the address; haversine is the metric. Inside a city the hex is exact. Between cities it is wrong — measured SF→Tokyo at 2.8× too far — so cross-city work uses haversineKm. The test suite asserts that distrust rather than the accuracy.
A new city is a new frame value, not a new codebase. Every function takes a CityFrame parameter.
npm test # 19 passing
Built by scripts/build-geohex.ts from Overpass, and the technique is ported from nyc-buildings — specifically the shape of its pipeline, not its data:
| Ported | Not ported |
|---|---|
| earcut triangulation per footprint | its regl + raw GLSL WebGL-only renderer |
one flat Float32Array for the whole city |
its 816 MB of pre-baked Manhattan binaries |
| one draw call, not one per building | |
| barycentric coords for edge treatment |
We build from live OSM so any city works.
Measured (--cell-km 0.5, bbox 30.235/-97.795 … 30.295/-97.705):
| OSM elements | 36,798 |
| buildings kept | 29,644 |
| triangles | 721,428 |
| unique hex cells | 108 |
| named | 1,438 (4.9%) |
| geometry served | 25,971,408 B — verified byte-for-byte in production |
Height provenance is kept per building, because a surveyed height and a guess are not the same claim:
height 26,535 ← the OSM height tag
levels 413 ← levels × 3.5m, an ESTIMATE
default 2,696 ← 8m fallback
The tallest-buildings table prints which is which on every row.
A guessed area threshold silently deleted 91% of the city. I set a 400 m² minimum; the script reported a clean run and produced 3,334 buildings. The real distribution over 36,584 Austin footprints is p10 18 m², p50 141 m², p90 377 m² — my cut sat above the 90th percentile. It is now the 25th percentile and a flag (--min-area), not a buried constant.
Every row printed [undefined] while the JSON manifest was correct: the loop iterated raw buildings (field heightSource) but read t.source. I nearly "fixed" the data instead of the printer.
Free text is judgement-bound to a cell. The hard part is not placing a post; it is deciding when you are allowed to say you know what it is.
Measured on a real listing:
"New Vietnamese spot opening next week on South Congress — bánh mì and cold brew, 12 seats, cash only."
has_specific_street 0.86 "yes, it names a street"
street_confidence 0.31 "...but I cannot place it"
Two vague questions about one fact, giving two vague answers — and the verdict depended on which one the threshold happened to read. That post lost its cell entirely.
Replaced with one question: pick the place from a closed list we supply. The model cannot invent a landmark, because the list is the only set of options it may return, and NONE is a real answer rather than a hedge:
place = South Congress (p = 1.00)
| Verdict | Meaning | Count |
|---|---|---|
identified |
placed and named → durable entity UUID | 2 |
anchored |
placed, not named → cell only, no identity | 2 |
unanchored |
kept, with the numbers that said so | 7 |
A cell can be located long before an entity can be identified. On the listing above:
place South Congress p=1.00 ← we can place it
named_properly 0.02 ← we CANNOT identify it
We mint a cell and refuse an identity. Guessing a name produces an entity that can never be deduplicated when the real listing arrives two posts later — it manufactures permanent garbage.
Nothing is dropped. A rant, a joke, an unplaceable post all return an observation carrying the raw probabilities. A missing answer records null, never a fake 0.
'Deceptively Good' 3a0b80ba-…-fdcddeefe61e cell 3403e7fe-0005-… @ The Domain
'The Basement' 3a0b80ba-…-fca2fad54c90 cell 3403e7fe-0005-… @ The Domain
Two named openings at one landmark: same cell, distinct entities. That is the addressing working end to end.
Every post is classified by jev-latest on the server. The key never reaches the browser — a feed that ships its own judgement key is a feed anyone can drain.
GET /api/observations classify the corpus live
POST /api/observations classify something you just submitted
Measured in production on austinsays.com:
HTTP 200 in 2.14s model jev-1.13.0
count 11 · identified 2 · anchored 2 · unanchored 7
durable UUIDs 2 · input tokens 20,748 (avg 1,886/post)
One request carries every question, because the questions are independent and run in parallel.
Writable, and tested with real input. POST of "New ramen bar on East 6th, 14 seats" returned ANCHORED at Sixth Street (p=0.77) with no entity UUID, because named_properly came back 0.03. The gate refusing an identity on live user input is better evidence than a passing test.
An error is not an empty feed. If Jev is unreachable the route returns 502 and the panel says "This is not an empty city; it is a city we could not read." An empty feed reads as "the city is quiet." Those are different claims.
Provenance is carried to the UI. Every post is tagged shit-says or written. The three shit.says posts are lifted verbatim from shit-says-feed/app/api/feed/[city]/route.ts; the rest were written to exercise the classifier, and are labelled as mine — a synthetic post presented as a real one is the exact dishonesty this module exists to prevent.
There is a second hex engine in the estate at outputs/dominos/src/hex — real, deployed at hxphxh.vercel.app/hex/td, with rings, spirals, A* pathfinding, pan/zoom hooks and a 3D instanced board. It is obviously worth reusing.
It is flat-top. We are pointy-top. The same six neighbour deltas, in a different ORDER — so a direction index means a different compass bearing in each. Port hexNeighbors() across and nothing throws. Neighbours come out mirrored, ring(centre, 1) walks the wrong way, and a UUID addresses the far side of the city from the building it describes.
That is the expensive kind of bug: silent, and in the addressing layer everything else trusts. So the convention is pinned by value in lib/hex-convention-guard.test.ts.
The test earned its keep on my own mistakes. I hand-wrote the direction table from memory and got two of six wrong. Deriving it correctly took three more tries:
- a 0.4 km probe is smaller than one 2 km cell → all six probes returned
(0,0) - 0.9 × cell put one diagonal back inside the origin cell (5 distinct, not 6)
- 1.2 × cell still collapsed six compass probes onto four cells, because in a pointy-top grid the neighbours are not at N/NE/SE/S/SW/NW
- and I asserted
−rwas north. Measured,+ris north —planeToLonLatmapsystraight to latitude andygrows upward. The test was failing on correct code.
deriveDirections() now throws rather than returning a wrong table.
9 of 11 posts kept but unidentified. That is the cost of only minting what we can evidence, and it is printed on the page rather than buried. A system reporting 11/11 would be lying.
80% of buildings carry a surveyed height; the rest are estimates and are labelled.
Jev costs $0.042/M input tokens, free output — a 6,000-token decision turn is ~$0.00025. Our feed: 20,748 tokens for 11 posts.
For calibration, we note what ClawPump's published numbers do when you weight agents equally rather than pooling records: input −37.3% → +55.4%. One agent was over half the traffic in both windows. Publish the number and the reason it doesn't hold.
- The name is still a regex. A post must clear
named_properlybefore we try, but once we try,extractName()reads a name out of free text. A model choosing from a bounded set could not manufacture a plausible wrong name. - One threshold bar, not two. Addressing a cell and minting an identity are different risks and should not share a number.
- Per-city data is inlined. The maths is multi-city; the venue tables are not yet a data module.
- The map does not visually distinguish surveyed from estimated heights, though the tables do.
- WFC is unused. Agreed as the v2 move for venues, districts, SpaceX, SXSW. The reference repo is WebGPU-only and renders a blank page on any device without it.
The build failed for a long time on Cannot resolve '@/lib/uuidv8-client' — a module that existed on disk and was tracked in git.
NODE_ENV=production # in the shellnpm honours it and omits devDependencies, so all ten were missing — including tailwindcss. A broken PostCSS pipeline makes Next report CSS-adjacent module resolution as "cannot resolve" for perfectly good TypeScript. The missing tsx was why npm test had been failing too.
npm install --include=devCheck the config before you start deleting the files the error names.
npm test # 19 passing
npx tsx scripts/build-geohex.ts --cache # rebuild the city from OSM
npx tsx scripts/jev-live-probe.ts # classify posts against live Jev
npm run dev # /city and /aboutshit.says—shit.<city>says.com, a parody character-tweet feed in the Overheard in Austin register. Already live atshit.sonomasays.com, and its city switcher already lists ten cities. The network shape exists; the per-city substance does not.- Next — Sonoma, on the wine. Then the WFC pass for venues and districts.
Data © OpenStreetMap contributors (ODbL). Optician Sans under SIL OFL.