Entur's geocoder: autocomplete and reverse geocoding for Norwegian stop places, addresses, place names and points of interest. Two components - a patched Photon search backend on OpenSearch, and a Ktor proxy in front of it serving a v3 API plus a Pelias-compatible v2. The search index is built from a Nominatim NDJSON dump produced by nominatim-converter.
flowchart LR
A[nominatim-converter] -->|nominatim.ndjson| B[Photon index]
B --> C[geocoder-photon]
C --> D[geocoder-proxy]
D --> E[api.entur.io/geocoder]
Public API: developer.entur.no/apis/geocoder -
/v3/autocomplete, /v3/reverse, /v3/place.
curl "https://api.entur.io/geocoder/v3/autocomplete?q=Oslo+S"One command converts every source with converter-prod.json, builds the index and starts
Photon. It downloads the whole country, so expect it to take a while.
cd photon
./import/download-photon-jar.sh
./full-local-reimport-and-start.sh
# In another terminal, from the repo root - or run no.entur.geocoder.proxy.AppKt from your IDE
./gradlew :proxy:runStarted from a console the proxy talks to http://localhost:2322; set PHOTON_URL to point it
somewhere else.
The same steps one at a time, for a different config, or to skip the conversion and use what CI built:
cd photon
./import/download-photon-jar.sh
# EITHER convert the data - import/config/converter-{prod,dev,local,sweden-test,denmark-test}.json
./import/create-nominatim-data.sh import/config/converter-local.json -z
# OR take the ndjson CI last built
./download-latest-nominatim-data.sh
./import/create-photon-data.sh nominatim.ndjson.gz
# OR skip both and take CI's finished index
rm -rf photon_data && ./download-latest-photon-data.sh
./photon-start.shcurl -s 'http://localhost:8080/v3/autocomplete?q=sk%C3%B8yen%20stasjon&limit=20'
curl -s 'http://localhost:8080/v3/reverse?lat=59.92&lon=10.67&radius=1&limit=10&layers=address,locality'
# v2 (Pelias-compatible) uses different parameter names
curl -s 'http://localhost:8080/v2/autocomplete?text=sk%C3%B8yen%20stasjon&size=20'&debug=true also reveals the native Photon results with importance (input weight) and
score (weight calculated by Photon).
Photon directly - category values are case-sensitive, so layer.stopPlace matches while
layer.stopplace silently returns nothing:
curl -s 'http://localhost:2322/api?q=Berglyveien&include=layer.stopPlace'OpenSearch directly. Document ids are the entity id with : replaced by -
(NSR-StopPlace-58404, KVE-PostalAddress-12191345), not numeric OSM ids - osm_id is 0 for
NSR documents, so use extra.id from the API response as the key.
curl -s 'http://localhost:9201/photon/_mapping' | jq . # available fields
curl -s 'http://localhost:9201/photon/_doc/NSR-StopPlace-58404' | jq .Handy for checking what the converter actually wrote, e.g. the alt names a multimodal parent inherited from its children:
$ curl -s 'http://localhost:9201/photon/_doc/NSR-StopPlace-58404' | jq -c '._source.name'
{"default":"Nationaltheatret","alt":"Nationaltheatret stasjon;Nasjonalteatret;Nationaltheatret"}The same works against a pod in GKE via a port-forward:
kubectl --context dev port-forward <geocoder-photon-pod> -n geocoder 9201
ID=$(curl -s 'https://geocoder-photon.dev.entur.io/api?q=ullerud' \
| jq -r '.features[0].properties.extra.id' | tr ':' '-') # NSR-StopPlace-5496
curl -s "http://localhost:9201/photon/_doc/$ID" | jq -c "[._source.importance, ._source.name.default]"
[0.078586,"Ullerud"]importance is set by the converter in the Nominatim data; score is what Photon computes
from it at query time.
$ curl -s 'http://localhost:8080/v2/autocomplete?text=Oslo&debug=true&size=1' \
| jq -c '.geocoding.debug.raw_data[] | [.localeTags.name.default, .infos.importance, .score]'
["Oslo",0.92,2.8717440524466067]
["Oslo S",0.538821,2.323909030548797]
["Oslo",0.27596,2.27596]
["Oslo lufthavn",0.550358,2.006965529061908](Debug shows three more results than asked for, see PhotonAutocompleteRequest.RESULT_PRUNING_HEADROOM. Both numbers change with every index build, so treat them as illustrative. The two "Oslo" rows are the group of stop places and the locality.)
All deployment runs from main. The daily import uses the prod-approved tag - remember to
move it when a commit is ready for production:
git tag -f prod-approved [sha]
git push origin prod-approved --force
| Workflow | Trigger | What it does |
|---|---|---|
| proxy.yml | push to main, manual |
Builds and deploys the proxy to dev; tst and prd need approval. Manual dispatch takes a target (dev only | dev → tst → prd | tst → prd) |
| proxy-deploy.yml | manual | Deploys an existing proxy image tag |
| photon-scheduled.yml | daily 06:27 UTC | Full import + build + deploy to tst → prd, no approval gates. Checks out prod-approved and updates the latest-prod.txt pointer |
| photon.yml | manual | Import, build image, deploy (same targets; optional config, default converter-prod.json) |
| photon-deploy.yml | manual | Deploys an existing Photon image tag |
All builds run acceptance tests after deployment, and most workflows post to Slack on failure. The reusable _generate-tag.yml and _deploy-and-test.yml workflows back the build and deploy jobs; shared steps live as composite actions under .github/actions/.
photon-sweden-scheduled.yml
runs a full Swedish import and deploy to dev every Monday at 05:27 UTC. It tracks main -
Sweden never reaches prod, so there is no prod-approved tag - and updates latest.txt. It
also keeps photon-data-se/ inside the bucket's 90-day lifecycle window, so a running pod's
photon_data.tar.gz can't be deleted out from under it. The manual counterparts are
photon-sweden.yml and
proxy-sweden.yml.
Denmark has manual-only equivalents,
photon-denmark.yml and
proxy-denmark.yml.
- cache-data-sources.yml - daily 03:00 UTC: downloads the third-party sources (matrikkel, stedsnavn, custom POIs from poiman) plus PostHog popular-stops, verifies size, and uploads them to
gs://ent-geocoder-prd/data-sources/. The nightly import reads from this cache rather than hitting upstream directly. - monitor-photon-data.yml - daily 08:22 UTC: checks
photonImportDatefrom the prod/v2/infoendpoint and alerts Slack if the data is older than 50h. - api-docs.yml - lints both OpenAPI specs on every push/PR touching
proxy/docs/**,openapi3.ymlor.spectral.yml; on push tomainpublishes the v3 spec to developer.entur.no/apis/geocoder andproxy/docs/to the docs portal. The v2 spec is linted but no longer published.
Built artifacts live in the public bucket gs://ent-geocoder-prd/:
| Prefix | Contents |
|---|---|
nominatim-data/ |
nominatim.ndjson.gz per build (+ .sha256) |
nominatim-data-se/ |
Sweden variant |
photon-data/ |
photon_data.tar.gz per build (+ .sha256) |
photon-data-se/ |
Sweden variant |
data-sources/ |
Daily-refreshed source files |
Each build writes to <prefix>/<tag>/<filename>. The <tag> is generated once and shared
between the docker image and the GCS upload, so geocoder-photon:<tag> always pairs with
gs://.../photon-data/<tag>/photon_data.tar.gz. Two pointer files at the prefix root track
recent builds: latest.txt (most recent build from any branch) and latest-prod.txt (most
recent build deployed to prod, written by photon-scheduled.yml).
The photon container fetches photon_data.tar.gz from $PHOTON_DATA_URL on startup, verifies
its .sha256 sidecar, and writes a photon_data/.ready sentinel after extraction so in-place
restarts skip the download. CI derives the URL from the image tag in
_deploy-and-test.yml and injects it into the helm
values; templates/photon-data-validation.yaml fails the render if it is missing.
# See the current pointer
curl -s https://storage.googleapis.com/ent-geocoder-prd/photon-data/latest-prod.txt
# Re-deploy a known-good image - the data is paired automatically
gh workflow run photon-deploy.yml -f target='tst → prd' -f image_tag=<previous-tag>Applied once per bucket. The matchesSuffix filter spares the latest*.txt pointer files.
{
"lifecycle": {
"rule": [
{
"action": {"type": "Delete"},
"condition": {
"age": 90,
"matchesPrefix": ["nominatim-data", "photon-data"],
"matchesSuffix": [".gz", ".sha256"]
}
}
]
}
}Ranking happens inside Photon, so most search tuning lands in the fork rather than here.
- Make the change in a checkout of komoot/photon and build it with
./gradlew build. - Create a tag and push it to entur/photon with
git push --tags entur. - Draft a release at entur/photon/releases/new, select the tag, attach
photon-<tag>.jarfrom Photon'starget/, check "Set as a pre-release" and publish. - Copy the asset link and update
PHOTON_JARin photon/import/download-photon-jar.sh. - Push, then run photon.yml with target
dev only.
Dashboards - Photon metrics · Proxy metrics
Ours
- nominatim-converter - builds the index this proxy queries. Importance lives in
src/common/importance.rs, categories insrc/common/category.rs - entur/photon - the patched fork that does the actual ranking. PhotonDocSerializer decides which name fields get indexed and at what priority
- geocoder-acceptance-tests
- bau - v2 vs v3 comparison tool (hosted)
External
- photon and the photon pelias adapter
- OSM dumps for photon from graphhopper
- Nominatim and its database layout