← Glossary Desk / API
Tokens

Drive Glossary Desk from your own code

Everything the web page does is available over HTTP: post one document in, get back either its domain glossary — canonical terms, tight definitions, the synonyms to avoid, the places the text contradicts itself, a ready CONTEXT.md — or a discovery questionnaire aimed at the one person who can settle what the document cannot. The natural use is a CI job that re-extracts the glossary whenever a spec changes and fails the build when the verdict comes back contradictory, or a batch that walks a folder of PRDs and reports which ones use four words for the same thing.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{ "ok": true,  "data":  { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "status": 402, "details": { ... } } }

Send your token as Authorization: Bearer … on every call. The app slug travels in the body of /guest as {"slug": "glossary-desk"}; after that the token itself carries the app, so a run needs only Authorization, Content-Type: application/json and the Idempotency-Key described in step 5.

The request body for /estimate, /run and /run-stream is the input object itself — not wrapped in anything. A body of {"input": {…}} returns 200 and quietly hides every field from the model, so guard it: the app's own client refuses to send anything that is not a plain JSON object.

Error codes

codestatuswhat to do
unauthorized401The token is missing, malformed or expired. Get a new one from the token page.
payment_required402The balance is below min_credits. Call /estimate first and top up.
forbidden403The token is valid but not for this app, or a guest token tried a metered run. Mint the guest token against glossary-desk, or sign in.
not_found404Unknown job id, unknown collection, or the app slug does not exist.
conflict409The same Idempotency-Key was replayed with a different body. Change the key or send the original input.
validation_error422The input is not a plain JSON object, or a field is the wrong type — needs must be an array of objects, prescan_facts an object. A body that is not valid JSON at all comes back as a 400.
rate_limited429Too many requests. Back off and retry; do not tight-loop.
internal5xxA server-side failure, reported as server_error on a plain 500. Retry with the SAME Idempotency-Key so you are not billed twice.

The field to get right first: task

Glossary Desk is one app with two lanes, and task is what chooses between them. It is the first field of every request body:

taskwhat comes backextra input fields
"glossary"The document's ubiquitous language: terms, ambiguities, excluded, a verdict of consistent, drifting or contradictory, and a context_md you can commit as CONTEXT.md.none
"questionnaire"A discovery questionnaire for one named recipient: purpose, from_line, to_line, use_line, themed questions most-important-first, a coverage table, and a questionnaire_md.recipient, needs, deadline

The system prompt routes on task and never blends the two contracts in one reply. A missing or unknown task is not an error: the model picks the closest lane from the fields that are present — a recipient or a needs array means questionnaire, anything else means glossary — names the lane it chose in the reply's lane field, and says so in notes_on_input. The client does the same thing defensively: a reply whose lane is neither id is read as questionnaire when it carries themes or questionnaire_md, and as glossary otherwise, with a sentence appended to notes_on_input saying the reply did not name its lane. So the degradation is always visible — but branch on lane in the reply, never on the task you think you sent.

The two lanes chain. Run glossary, take the ambiguities whose decidable_from_text is false, feed their text in as the needs of a questionnaire run, collect the answers, and paste them back into a second glossary run as decisions — which are treated as settled and outrank the document.

1. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool — it reads the same storage the app itself uses and prints the token for you.

A guest token can call /me and /estimate. Extracting a glossary or writing a questionnaire is metered, so it needs a personal token from signing in.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://glossary-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="aut_YOUR_TOKEN"
#
# To mint a guest token from the command line instead. A guest token is enough for
# /me and /estimate; a glossary or a questionnaire run needs a personal token.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" \
  -d '{"slug": "glossary-desk"}'
# {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}

2. A tiny client

One helper that adds the headers, unwraps data and raises on error.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="glossary-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://glossary-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

3. Check the session and the balance

GET /me tells you whether the token is a guest or a person, and what the balance is. The object is small and carries exactly three things: subject_type — guest or user, and a guest can price a run but not start one — subject_id, and credits, the wallet balance. There is no username in it, so "signed in" is subject_type == "user" and nothing else. Compare credits against min_credits from the next step before you run, so a shortfall surfaces as your own clear message rather than a 402.

call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}

4. Price the run — free

The input object is exactly what the app's own form submits. It is always a JSON object — never a bare string, never wrapped in an input key:

fieldtypemeaning
taskstring"glossary" or "questionnaire". The lane. Missing or unknown degrades to the closest lane by the fields present, and the reply names what it chose in lane and in notes_on_input.
documentstring, requiredThe pasted text: a spec, a PRD, a design note, a meeting transcript, a support thread, a policy, an RFC. This is the run's only evidence — every term, definition, ambiguity and question has to rest on words that are in it. The browser normalises line endings and clips to 60,000 characters from the middle, keeping the beginning and the end, and leaves the marker [... N characters cut from the middle of the document - the beginning and the end are kept ...] in place of the cut. Clip the same way if you send more, and keep the marker: the prompt keys on it, refuses to claim anything about the missing middle, and mentions the cut in notes_on_input.
context_hintstring, optionalOne or two sentences about what the project is and who wrote the document — "A B2B SaaS product. The PRD was written by a PM and edited by two engineers." The browser caps it at 2,000 characters.
decisionsstring, optionalSettled answers from an earlier questionnaire, as free text. Authoritative: a term a decision settles comes back firm with the decision as its evidence, an ambiguity the decisions resolve is not raised again, and a decision that contradicts the document still wins — with the contradiction noted in notes_on_input. Capped at 8,000 characters.
prescan_factsobjectWhat a local scanner found, for free, before the run. Shape below. Empty arrays are legitimate — read the honesty note. The platform's input check declares scalar fields only, so /estimate answers an object-valued prescan_facts or needs with an advisory unknown field warning; that warning is expected and the field still reaches the model. Only task and document are required.
retry_notestring, optionalSend only on a retry, when a previous reply failed to parse or came back truncated. The instruction is obeyed exactly. The web app adds it automatically on its one automatic retry.
questionnaire lane only
recipientstringOne named person, their role, and how you relate to them — "Maya, Head of Billing. She owns the pricing and refund rules and signs off invoicing changes. I am the PM on the plan-change feature." Without it the questionnaire is pitched at an unnamed colleague, which is weaker. Capped at 2,000 characters.
needsobject[][{"id": "N-001", "text": "…"}] — what you need back, one decision per entry, ids sequential from N-001. The browser numbers up to 20 lines of typed text (6,000 characters) and strips any bullet or number the user typed. Every id you send comes back exactly once in coverage. Omitting needs is allowed: the model derives them from the document's own open items and says so, and coverage comes back empty.
deadlinestring, optionalWhen you need the answers and how long it should take — "Answers by Thursday; 20 minutes should be enough." Capped at 2,000 characters.

prescan_facts, honestly

In the browser this object is computed for free by a local scanner before the run:

{
  "stats": { "words": 210, "sentences": 14, "headings": 5, "paragraphs": 9 },
  "candidates": [
    { "id": "C-001", "term": "customer", "count": 4,
      "forms": ["customer", "customers"], "first_seen": "line 5", "strong": true }
  ],
  "clusters": [
    { "id": "K-001", "terms": ["customer", "user", "client", "account"],
      "reason": "known synonym family, all four appear" }
  ],
  "redefinitions": [
    { "id": "R-001", "term": "cancellation", "kind": "conflicting_rules",
      "quotes": ["Cancellation takes effect immediately", "keeps read-only access until the end of the period"] }
  ],
  "clipped": { "cut": 0 }
}

candidates are the repeated noun phrases the scanner thinks might be terms, with their occurrence count and surface forms; clusters are groups of words that look like competing names for one thing; redefinitions are places the same word is defined twice or where two rules about it conflict — kind is defined_twice or conflicting_rules. clipped.cut is how many characters the middle clip removed.

An API caller does not have to reproduce any of that. Sending {"stats": {...}, "candidates": [], "clusters": [], "redefinitions": [], "clipped": {"cut": 0}} is legitimate and the run still works — the model reads document either way. Be honest with yourself about what you give up: the glossary lane's coverage list comes back empty, because it is one entry per candidate id, and the client-side reconciliation — which checks that every candidate came back with a status, that every cluster and redefinition was addressed somewhere, and that no id you did not send appears — then has nothing to check. You lose the accountability, not the glossary. The questionnaire lane is unaffected: its coverage is keyed on needs, not on candidates, so empty candidate arrays cost it nothing.

What makes the facts worth sending is that contract: every id in prescan_facts.candidates comes back exactly once in coverage, with a status of defined, avoid, excluded, ambiguity or not_a_term. A mechanical scanner is allowed to be wrong, and not_a_term is the honest answer for a phrase that only looked like a term — that is a different thing from silence. A candidate that never appears at all is a failed run, not a passing one.

/estimate creates no job and charges nothing. It returns the model binding — model is gpt-5.6-terra, model_alias is gpt-terra, markup_bps is 1000 — and the reservation: hold_credits is what gets held, and min_credits is the balance you must clear to start at all. The hold is a reservation, not the price. It prices the full output cap, so the charged_credits you see on the settled job is usually far lower. Budget against hold_credits, report against charged_credits.

Price each lane separately. The questionnaire body is shorter than a glossary body, and the same document with an empty prescan_facts prices differently from one carrying forty candidates — the facts are input tokens like everything else.

# Lane A - the glossary. One document, the facts a local scan found, no decisions yet.
INPUT='{"task": "glossary", "document": "# Subscription billing - PRD v0.3\n\n## Plans\nWe sell three tiers: Starter, Team and Business. A user picks a package at sign-up and can change tier at any time from the billing page.\n\n## Cancellation\nA client can cancel from the billing page. Cancellation takes effect immediately: access ends and no further invoices are raised.\n\n## Invoices\nEvery charge produces an invoice, emailed to the billing contact. Failed payments retry three times over seven days, after which the subscription is cancelled - the customer keeps read-only access until the end of the period they already paid for.\n\n## Seats\nTeam and Business are priced per seat. A seat is an invited member who has accepted; pending invitations do not count.", "context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.", "decisions": "", "prescan_facts": {"stats": {"words": 210, "sentences": 14, "headings": 5, "paragraphs": 9}, "candidates": [{"id": "C-001", "term": "customer", "count": 4, "forms": ["customer", "customers"], "first_seen": "line 14", "strong": true}, {"id": "C-002", "term": "seat", "count": 3, "forms": ["seat", "seats"], "first_seen": "line 17", "strong": true}, {"id": "C-003", "term": "billing page", "count": 3, "forms": ["billing page"], "first_seen": "line 5", "strong": false}], "clusters": [{"id": "K-001", "terms": ["customer", "user", "client"], "reason": "known synonym family, all three appear"}], "redefinitions": [{"id": "R-001", "term": "cancellation", "kind": "conflicting_rules", "quotes": ["Cancellation takes effect immediately: access ends", "the customer keeps read-only access until the end of the period they already paid for"]}], "clipped": {"cut": 0}}}'

call estimate "$INPUT"
# {"ok":true,"data":{"model":"gpt-5.6-terra","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":2652,"min_credits":310}}
#
# estimate is FREE. It creates no job and charges nothing. hold_credits is what
# gets RESERVED; charged_credits on the settled job is normally much lower.

# Lane B - the questionnaire over the same document, aimed at one person, with
# two things only she can settle. Empty prescan arrays are fine here: this lane's
# coverage is keyed on `needs`, not on candidates.
Q_INPUT='{"task": "questionnaire", "document": "# Subscription billing - PRD v0.3\n\n## Cancellation\nA client can cancel from the billing page. Cancellation takes effect immediately: access ends and no further invoices are raised.\n\n## Invoices\nFailed payments retry three times over seven days, after which the subscription is cancelled - the customer keeps read-only access until the end of the period they already paid for.", "context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.", "decisions": "", "recipient": "Maya, Head of Billing. She owns the pricing and refund rules and signs off invoicing changes. I am the PM on the plan-change feature.", "needs": [{"id": "N-001", "text": "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both."}, {"id": "N-002", "text": "Whether customer, user and client are one thing or several, and which word Billing uses on invoices."}], "deadline": "Answers by Thursday; 20 minutes should be enough.", "prescan_facts": {"stats": {"words": 95, "sentences": 5, "headings": 3, "paragraphs": 4}, "candidates": [], "clusters": [], "redefinitions": [], "clipped": {"cut": 0}}}'

call estimate "$Q_INPUT"

5. Run it, then poll

POST /run returns a job_id; poll GET jobs/{job_id} until status is succeeded or failed. The result JSON is the string at data.output.output. The terminal job also carries charged_credits — the real price — and the truncated flag.

Always send an Idempotency-Key. The web app builds it as glossary-desk:<task>:<input hash>:a<attempt> and so should you. Three parts, three reasons:

A retried request carrying the same key returns the same job instead of billing a second run, which is what makes a CI retry safe after a network blip.

# The key is slug:task:hash:attempt. A retried request with the same key returns
# the SAME job instead of billing a second run.
KEY="glossary-desk:glossary:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

# Poll until the job reaches a terminal status.
while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# The terminal job looks like this:
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"glossary\",\"title\":\"Subscription billing\", ...}"},
#   "charged_credits":588,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])'

6. Or stream it

POST /run-stream is the same call over server-sent events, and it takes the same Idempotency-Key. Each delta event carries {"text": "..."}, a chunk of the result JSON, and the final done event carries status, charged_credits — the real price, normally a fraction of the hold — and the truncated flag.

Read the SSE yourself. What a client receives depends on where it is: a command-line reader like the ones below gets real delta events, while the same endpoint sends a page in a browser tick heartbeats instead — so a JavaScript callback wired to deltas never fires there, and any progress display, streaming preview or partial-recovery path built on it is dead code in a browser. Parse the event stream in your own reader, as the samples here do, and treat a run with no deltas at all as normal rather than as a stall: wait for done, or fall back to /run and polling.

The practical tip: the web app does not parse the partial JSON to drive its progress display, it watches for key names arriving in the accumulating text. In the glossary lane the appearance of "terms", then "ambiguities", then "coverage", then "context_md" is what advances the stage from choosing the canonical terms to quoting the evidence, reconciling the scan, and finally writing the CONTEXT.md. In the questionnaire lane the sequence is "purpose", "context_paragraph", "themes", "coverage", "questionnaire_md". Substring matching on the quoted key name is enough, and it costs nothing.

# Server-sent events. From the command line each `delta` carries a chunk of the
# JSON; the final `done` event carries the status, charged_credits and truncated.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"glossary\",\"title\":\"Subscription"}
# event: delta  {"text":" billing - PRD v0.3\",\"verdict\":\"contradictory\","}
# event: done   {"status":"succeeded","charged_credits":588,"truncated":false}
#
# A page in a browser gets `tick` heartbeats here instead of `delta` events, so
# read the stream in your own reader rather than relying on a delta callback.

7. Parse the result

data.output.output is a string holding one JSON object — unwrap twice. The web app strips an optional code fence, takes everything from the first { to the last }, parses that, and then normalizes it. Doing the same two things — the slice and the normalization — is what makes a caller robust against the small variations a model produces.

Here is an abbreviated glossary reply for the billing PRD above, structurally complete:

{
  "lane": "glossary",
  "title": "Subscription billing - PRD v0.3",
  "summary": "The PRD describes plan changes, trials, cancellation, pausing and per-seat pricing for a three-tier B2B subscription. The one thing a reader must know is that cancellation is stated two incompatible ways, and every other rule about access depends on which one is true.",
  "context_name": "Subscription Billing",
  "context_summary": "The context that owns plans, subscriptions, invoices and seats for a paying organisation.",
  "verdict": "contradictory",
  "verdict_reason": "Cancellation is defined twice with opposite consequences for access, so no rule about end-of-access can be implemented from this document.",
  "terms": [
    { "id": "T-001", "term": "Subscription", "definition": "One organisation's paid relationship to a tier, billed per period and per seat.",
      "avoid": ["package"], "group": "Core", "evidence": "A user picks a package at sign-up and can change tier at any time",
      "confidence": "firm" },
    { "id": "T-002", "term": "Seat", "definition": "An invited member who has accepted the invitation; the unit Team and Business are priced by.",
      "avoid": [], "group": "Pricing", "evidence": "A seat is an invited member who has accepted; pending invitations do not count",
      "confidence": "firm" },
    { "id": "T-003", "term": "Customer", "definition": "The organisation that holds the subscription and receives the invoice.",
      "avoid": ["client", "user"], "group": "Core", "evidence": "A client can cancel from the billing page",
      "confidence": "tentative" }
  ],
  "ambiguities": [
    { "id": "A-001", "term": "Cancellation", "kind": "contradiction",
      "summary": "Cancellation ends access immediately in one section and at the end of the paid period in another.",
      "evidence": ["Cancellation takes effect immediately: access ends and no further invoices are raised",
                   "the customer keeps read-only access until the end of the period they already paid for"],
      "proposal": "Name the two cases separately: voluntary cancellation ends access at once; cancellation after failed payment leaves read-only access until the paid period ends.",
      "decidable_from_text": false, "owner_hint": "Head of Billing" },
    { "id": "A-002", "term": "Customer", "kind": "synonyms",
      "summary": "customer, user and client are used interchangeably for the party that holds the subscription.",
      "evidence": ["A user picks a package at sign-up", "A client can cancel from the billing page"],
      "proposal": "Use Customer for the paying organisation and Member for a person inside it.",
      "decidable_from_text": true, "owner_hint": "" }
  ],
  "excluded": [
    { "term": "billing page", "reason": "A screen in the product, not a concept of the billing domain." }
  ],
  "coverage": [
    { "id": "C-001", "term": "customer",     "status": "defined",   "ref": "T-003" },
    { "id": "C-002", "term": "seat",         "status": "defined",   "ref": "T-002" },
    { "id": "C-003", "term": "billing page", "status": "excluded",  "ref": "" }
  ],
  "credential_seen": false,
  "notes_on_input": "",
  "context_md": "# Subscription Billing\n\nThe context that owns plans, subscriptions, invoices and seats...\n\n## Terms\n\n**Subscription**: One organisation's paid relationship to a tier...\n\n**Seat**: An invited member who has accepted...\n\n**Customer**: The organisation that holds the subscription...\n\n## Avoid\n\n- package (use Subscription)\n- client, user (use Customer)\n"
}

And the questionnaire reply for the same document, aimed at Maya:

{
  "lane": "questionnaire",
  "title": "Questions for the Head of Billing",
  "summary": "Two decisions block the plan-change feature: what cancellation does to access, and which word names the paying party.",
  "purpose": "Settle the cancellation rule and the naming of the paying party so the PRD can be implemented.",
  "from_line": "From: the PM on the plan-change feature",
  "to_line": "To: Maya, Head of Billing",
  "use_line": "Your answers go straight into PRD v0.4 and into the glossary the team builds from.",
  "context_paragraph": "The PRD lets a customer change tier without a support ticket. Two rules in it point in opposite directions, and both are yours to decide.",
  "how_to_answer": "Answer under each question. One or two sentences is plenty; where you are unsure, say so and name who is.",
  "themes": [
    { "heading": "Cancellation and access",
      "questions": [
        { "id": "Q-001", "question": "When a customer cancels voluntarily, does access end that moment or at the end of the period they have paid for?",
          "why": "The PRD states both, and every downstream rule about invoices and read-only access depends on the answer.",
          "covers": ["N-001"], "priority": 1 },
        { "id": "Q-002", "question": "Should cancellation after three failed payments behave differently from a cancellation the customer chose?",
          "why": "The failed-payment section grants read-only access that the voluntary path denies.",
          "covers": ["N-001"], "priority": 1 }
      ] },
    { "heading": "What we call the paying party",
      "questions": [
        { "id": "Q-003", "question": "Which single word should invoices and the billing page use for the party that pays: customer, client or account?",
          "why": "Three words appear in one document, and the invoice template has to pick one.",
          "covers": ["N-002"], "priority": 2 }
      ] }
  ],
  "closing": "Anything else about billing that you expect to change this quarter and that I should not design around?",
  "coverage": [
    { "id": "N-001", "need": "Whether a cancellation ends access immediately or at the end of the paid period.",
      "status": "covered", "question_ids": ["Q-001", "Q-002"], "note": "Split into the voluntary and the failed-payment case." },
    { "id": "N-002", "need": "Whether customer, user and client are one thing or several.",
      "status": "partial", "question_ids": ["Q-003"],
      "note": "The question settles the invoice wording; whether user is a different concept is answerable from the document itself." }
  ],
  "credential_seen": false,
  "notes_on_input": "",
  "questionnaire_md": "# Questions for the Head of Billing\n\n**To:** Maya, Head of Billing\n\n## Cancellation and access\n\n### When a customer cancels voluntarily, does access end that moment or at the end of the period they have paid for?\n\n>\n\n"
}

What the normalizer does to it

The web app does not trust the reply verbatim, and neither should a caller. These are the behaviours you will actually hit:

The verdict rule

The glossary verdict is not free-form, and it is worth re-deriving rather than trusting:

contradictory  if any ambiguity has kind == "contradiction"
drifting       else if any ambiguity has kind "synonyms" or "overloaded"
consistent     otherwise

The client computes this from the ambiguity list it received and warns when the returned verdict disagrees, treating the ambiguity list as authoritative. Do the same: a gate that reads verdict alone can be talked out of failing by a reply that lists a contradiction and then calls itself drifting.

Invariants worth asserting in CI

# The result JSON is a string inside the envelope, so unwrap it twice.
RESULT=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])')

printf '%s' "$RESULT" | python3 -c '
import sys, json
r = json.load(sys.stdin)
print(r["lane"], "|", r["title"])
print(r["verdict"], "-", r["verdict_reason"])
for t in r["terms"]:
    avoid = ", ".join(t["avoid"]) or "-"
    print("  %s %-14s %-9s avoid: %s" % (t["id"], t["term"], t["confidence"], avoid))
for a in r["ambiguities"]:
    who = "the text settles it" if a["decidable_from_text"] else "needs " + (a["owner_hint"] or "a person")
    print("  %s %-14s %-14s %s" % (a["id"], a["term"], a["kind"], who))
'

# Re-derive the verdict from the ambiguities rather than trusting the field.
printf '%s' "$RESULT" | python3 -c '
import sys, json
r = json.load(sys.stdin)
kinds = [a["kind"] for a in r["ambiguities"]]
want = "contradictory" if "contradiction" in kinds else \
       "drifting" if ("synonyms" in kinds or "overloaded" in kinds) else "consistent"
if r["verdict"] != want:
    raise SystemExit("verdict says %s but the ambiguities make it %s" % (r["verdict"], want))
print("verdict reconciles:", want)
'

# Every prescan candidate id must come back exactly once in coverage.
printf '%s' "$RESULT" | python3 -c '
import sys, json
seen = [c["id"] for c in json.load(sys.stdin)["coverage"]]
want = ["C-001", "C-002", "C-003"]
bad = [i for i in want if seen.count(i) != 1] + [i for i in seen if i not in want]
if bad:
    raise SystemExit("coverage drift: " + ", ".join(bad))
print("coverage reconciles")
'

# The CONTEXT.md is ready to commit as-is.
printf '%s' "$RESULT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["context_md"])' > CONTEXT.md

The output contract

Every key in the object, as the web app reads it. First the envelope both lanes share:

keytypemeaning
laneenumglossary or questionnaire — which contract the rest of the object follows. Inferred from the fields present when the model omits it or invents a value, and the inference is recorded in notes_on_input. Branch on this, not on the task you sent.
titlestringA short name for this run, taken from the document's own subject. Empty becomes Untitled.
summarystringTwo to four sentences: what the document is about and the one thing the reader must know.
coverageobject[]The reconciliation table. Shape differs per lane — see below. One entry per prescan_facts.candidates id in the glossary lane, one per needs id in the questionnaire lane.
credential_seenbooleantrue when the paste looked like it carried a password, API key, token or private key. The reply then repeats no part of the value anywhere and says in notes_on_input that it should be rotated.
notes_on_inputstring"" when there is nothing to say. Carries: that the middle of the document was clipped, that a decisions entry contradicts the document, that the task was missing or unrecognised and which lane was chosen instead.

Glossary body

keytypemeaning
context_namestringThe bounded context the document is about, in two or three words — "Subscription Billing". Empty falls back to title.
context_summarystringOne or two sentences saying what this context owns and where its edge is.
verdictenumconsistent, drifting or contradictory. Anything else normalizes to drifting. The single value a CI gate should branch on — after re-deriving it from the ambiguity kinds.
verdict_reasonstringOne sentence naming the ambiguity that decided the verdict.
termsobject[]{id, term, definition, avoid, group, evidence, confidence}. Ids are sequential T-001, T-002, … definition is one or two sentences; avoid is the words the document uses for the same thing that the team should stop using; group is a heading such as Core or Pricing, defaulting to General; evidence is a verbatim quote from the document, trimmed to about 200 characters.
ambiguitiesobject[]{id, term, kind, summary, evidence, proposal, decidable_from_text, owner_hint}. Ids are sequential A-001, … evidence is an array of verbatim quotes; proposal is the resolution the model would pick; decidable_from_text says whether the document itself settles it, and when it is false, owner_hint names the kind of person who can. Those two fields are exactly what the questionnaire lane consumes.
excludedobject[]{term, reason} — words considered and deliberately left out, because they are general programming, business or project-management vocabulary the document merely uses rather than concepts of this context. Shipping the rejects is what makes the term list auditable.
coverageobject[]{id, term, status, ref}. One entry per prescan candidate id, exactly once, with status in defined, avoid, excluded, ambiguity, not_a_term, and ref pointing at the T- or A- id that handled it ("" otherwise).
context_mdstringA whole CONTEXT.md in one JSON string, ready to commit at the repository root: the context name, the summary, each term as **Term**: definition exactly once, and an Avoid section. "" when the document gives nothing to base it on.

Questionnaire body

keytypemeaning
purposestringOne sentence saying what the answers will be used for.
from_line / to_linestringWho is asking and who is being asked, drawn from recipient.
use_linestringWhat happens to the answers — the sentence that makes the recipient's time feel spent rather than taken.
context_paragraphstringEnough of the document for the recipient to answer without reading it.
how_to_answerstringThe instruction line: answer under each question, one or two sentences, say when you are unsure.
themesobject[]{heading, questions}, where each question is {id, question, why, covers, priority}. Question ids run Q-001 upward across themes, not per theme. why says what changes depending on the answer; covers lists the N- need ids the question serves; priority is 1, 2 or 3, most important first.
closingstringThe catch-all question at the end.
coverageobject[]{id, need, status, question_ids, note}. One entry per need you sent, exactly once. status is covered, partial, answered_in_document — the document already settles it, so no question was spent on it — or not_covered. question_ids names the questions that serve it.
questionnaire_mdstringThe whole questionnaire as Markdown in one JSON string, with each question as an ### heading and a blockquote answer stub — a line that is just > — under it, ready to paste into an email or a doc.

The enums

fieldvaluesnotes
verdictconsistent, drifting, contradictoryconsistent: the document names one thing one way. drifting: it has competing names or an overloaded word, but nothing that makes two rules incompatible. contradictory: at least one place where following the document two ways gives two different systems. Derived, not chosen — see the verdict rule. Unrecognised values normalize to drifting.
ambiguities[].kindoverloaded, synonyms, undefined, contradiction, boundaryoverloaded: one word carries two meanings. synonyms: several words carry one meaning. undefined: a word the document leans on but never defines. contradiction: two statements that cannot both hold. boundary: it is unclear whether a concept belongs to this context at all. Unrecognised values normalize to undefined.
terms[].confidencefirm, tentativefirm means the document (or an authoritative decisions entry) states the definition; tentative means it was inferred from usage and should be confirmed. Unrecognised values normalize to tentative, which is the safe direction.
coverage[].status (glossary)defined, avoid, excluded, ambiguity, not_a_termWhat became of a scanned candidate. not_a_term is a legitimate answer — the scanner is mechanical and is allowed to be wrong. Unrecognised values normalize to excluded.
coverage[].status (questionnaire)covered, partial, answered_in_document, not_coveredanswered_in_document means no question was spent on it because the text already settles it. Unrecognised values normalize to partial.
questions[].priority1, 2, 3A number, not a string. 1 is "the project stops without this". Anything else normalizes to 2.
redefinitions[].kind (input)defined_twice, conflicting_rulesSent by you inside prescan_facts, not returned.

The reply never echoes a secret value. If the paste contains an API key, a password in a connection string or a token in a log line, credential_seen comes back true and the note says to rotate it — the value itself appears in no term, no quote and no context_md.

8. Use it in CI

The worked example: a job reads the spec out of the repository, extracts its glossary, writes CONTEXT.md, and exits non-zero when the verdict is contradictory — the document says two incompatible things and nobody can implement it as written. Fail on contradictory; report on drifting, which is normal in a living document and would otherwise make the gate noise people learn to ignore. Derive the Idempotency-Key from the spec's contents so a re-run of the same commit replays the same job instead of re-billing, and only bump the attempt suffix when the text actually changed.

The script sends prescan_facts with empty arrays, which is the honest thing for a caller with no local scanner — and it means coverage comes back empty, so the gate leans on the verdict and the ambiguity list instead of on reconciliation. If you do have ids to send, send them: the reconciliation check is the strongest signal in the reply.

#!/bin/sh
# glossary-gate.sh - fail the build when the spec contradicts itself.
set -eu

BASE="https://api.skillsafe.ai/v1/app-api"
TOKEN="$SKILLSAFE_TOKEN"   # from https://glossary-desk.skillsafe.ai/tokens.html
SPEC="${1:-docs/spec.md}"

[ -f "$SPEC" ] || { echo "glossary-desk: no $SPEC in this repository"; exit 0; }

# 1. Build the input. An API caller may send empty prescan facts - and then the
#    coverage list comes back empty, so the gate reads the verdict instead.
INPUT=$(SPEC="$SPEC" python3 -c '
import json, os, pathlib
text = pathlib.Path(os.environ["SPEC"]).read_text(encoding="utf-8")
if len(text) > 60000:                       # clip the MIDDLE, keep both ends
    head, tail = text[:37000], text[-20000:]
    cut = len(text) - len(head) - len(tail)
    text = head + "\n\n[... %d characters cut from the middle of the document - the beginning and the end are kept ...]\n\n" % cut + tail
print(json.dumps({
    "task": "glossary",
    "document": text,
    "context_hint": "CI gate on every pull request that touches the spec.",
    "decisions": "",
    "prescan_facts": {"stats": {}, "candidates": [], "clusters": [],
                      "redefinitions": [], "clipped": {"cut": 0}},
}))')

KEY="glossary-desk:glossary:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(curl -sS "$BASE/jobs/$JOB" -H "Authorization: Bearer $TOKEN")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# 2. Gate on the verdict, re-derived from the ambiguities. Fail on
#    contradictory; report drifting without failing.
printf '%s' "$OUT" | python3 -c '
import sys, json, pathlib
job = json.load(sys.stdin)["data"]
if job.get("truncated"):
    raise SystemExit("::error::glossary-desk: reply was cut short, the glossary is incomplete")
r = json.loads(job["output"]["output"])
kinds = [a["kind"] for a in r["ambiguities"]]
verdict = "contradictory" if "contradiction" in kinds else \
          "drifting" if ("synonyms" in kinds or "overloaded" in kinds) else "consistent"
pathlib.Path("CONTEXT.md").write_text(r["context_md"], encoding="utf-8")
for a in r["ambiguities"]:
    who = "the text settles it" if a["decidable_from_text"] else (a["owner_hint"] or "a person")
    print("%s  %-14s %-14s %s  [%s]" % (a["id"], a["kind"], a["term"], a["summary"], who))
print(verdict, "-", r["verdict_reason"])
if verdict == "contradictory":
    blocking = [a["id"] for a in r["ambiguities"] if a["kind"] == "contradiction"]
    raise SystemExit("::error::glossary-desk: the spec contradicts itself (%s)" % ",".join(blocking))
if verdict == "drifting":
    print("::warning::glossary-desk: the spec is drifting - competing names, but nothing blocking")
'

Truncation and partial results

When the balance sits between min_credits and hold_credits, the run is not refused: it executes with a reduced output cap and comes back with truncated: true on the finished job and on the streaming done event. What you hold then is a prefix of the reply, not the reply — in the glossary lane the terms may be complete while coverage, excluded and context_md are missing or cut mid-string; in the questionnaire lane the themes may be there while the coverage table and questionnaire_md are not.

Check the flag before you treat a reply as complete, and remember that the client's parser is strict in one direction only: a prefix that still contains at least one term or one ambiguity parses, so a truncated glossary can look like a small glossary. That is why the flag, and not the shape of the object, is the test.

The right response is a retry, not a repair: resubmit with a retry_note asking for fewer, denser terms and a shorter context_md, and with the attempt suffix on the Idempotency-Key incremented so the new body is not a replay of the old key. Repairing truncated JSON by appending closing braces produces something that parses and is not what the model meant — and in this app it produces a glossary whose coverage silently disagrees with the terms above it.

One more honest limit: the document is clipped from the middle at 60,000 characters, and the reply will say so in notes_on_input. A glossary built from a clipped document is a glossary of the beginning and the end. For anything longer, split the document along its own section boundaries, run each part, and reconcile the term lists yourself — the terms carry evidence quotes precisely so that merging two runs is a matter of comparing quotes rather than trusting two definitions.