Drive Glossary Desk from your own code
Everything the web page does is available over HTTP: post one document in, get back either its
domain glossary — canonical terms, tight definitions, the synonyms to avoid, the places the text
contradicts itself, a ready CONTEXT.md — or a discovery questionnaire aimed at the one
person who can settle what the document cannot. The natural use is a CI job that re-extracts the
glossary whenever a spec changes and fails the build when the verdict comes back
contradictory, or a batch that walks a folder of PRDs and reports which ones use four
words for the same thing.
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{ "ok": true, "data": { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "status": 402, "details": { ... } } }
Send your token as Authorization: Bearer … on every call. The app slug travels in the
body of /guest as {"slug": "glossary-desk"}; after that the token
itself carries the app, so a run needs only Authorization,
Content-Type: application/json and the Idempotency-Key described in step 5.
The request body for /estimate, /run and /run-stream is the
input object itself — not wrapped in anything. A body of
{"input": {…}} returns 200 and quietly hides every field from the model, so guard it:
the app's own client refuses to send anything that is not a plain JSON object.
Error codes
| code | status | what to do |
|---|---|---|
unauthorized | 401 | The token is missing, malformed or expired. Get a new one from the token page. |
payment_required | 402 | The balance is below min_credits. Call /estimate first and top up. |
forbidden | 403 | The token is valid but not for this app, or a guest token tried a metered run. Mint the guest token against glossary-desk, or sign in. |
not_found | 404 | Unknown job id, unknown collection, or the app slug does not exist. |
conflict | 409 | The same Idempotency-Key was replayed with a different body. Change the key or send the original input. |
validation_error | 422 | The input is not a plain JSON object, or a field is the wrong type — needs must be an array of objects, prescan_facts an object. A body that is not valid JSON at all comes back as a 400. |
rate_limited | 429 | Too many requests. Back off and retry; do not tight-loop. |
internal | 5xx | A server-side failure, reported as server_error on a plain 500. Retry with the SAME Idempotency-Key so you are not billed twice. |
The field to get right first: task
Glossary Desk is one app with two lanes, and task is what chooses between them. It is
the first field of every request body:
task | what comes back | extra input fields |
|---|---|---|
"glossary" | The document's ubiquitous language: terms, ambiguities, excluded, a verdict of consistent, drifting or contradictory, and a context_md you can commit as CONTEXT.md. | none |
"questionnaire" | A discovery questionnaire for one named recipient: purpose, from_line, to_line, use_line, themed questions most-important-first, a coverage table, and a questionnaire_md. | recipient, needs, deadline |
The system prompt routes on task and never blends the two contracts in one reply. A
missing or unknown task is not an error: the model picks the closest
lane from the fields that are present — a recipient or a needs array
means questionnaire, anything else means glossary — names the lane it
chose in the reply's lane field, and says so in notes_on_input. The
client does the same thing defensively: a reply whose lane is neither id is read as
questionnaire when it carries themes or questionnaire_md,
and as glossary otherwise, with a sentence appended to notes_on_input
saying the reply did not name its lane. So the degradation is always visible — but branch on
lane in the reply, never on the task you think you sent.
The two lanes chain. Run glossary, take the ambiguities whose
decidable_from_text is false, feed their text in as the
needs of a questionnaire run, collect the answers, and paste them back
into a second glossary run as decisions — which are treated as settled
and outrank the document.
1. Get a token
The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool — it reads the same storage the app itself uses and prints the token for you.
A guest token can call /me and /estimate. Extracting a
glossary or writing a questionnaire is metered, so it needs a personal token from
signing in.
# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
# https://glossary-desk.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="aut_YOUR_TOKEN"
#
# To mint a guest token from the command line instead. A guest token is enough for
# /me and /estimate; a glossary or a questionnaire run needs a personal token.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" \
-d '{"slug": "glossary-desk"}'
# {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}
# Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot run a metered lane.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest",
data=json.dumps({"slug": "glossary-desk"}).encode(),
method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
// Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "glossary-desk" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest",
bytes.NewReader([]byte(`{"slug": "glossary-desk"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct {
Data struct {
Token string `json:"token"`
} `json:"data"`
}
_ = json.NewDecoder(guestRes.Body).Decode(&guest)
fmt.Println(guest.Data.Token)
// Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
var http = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\": \"glossary-desk\"}"))
.build();
HttpResponse<String> guest = http.send(guestReq, HttpResponse.BodyHandlers.ofString());
System.out.println(guest.body()); // {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}
# Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot run a metered lane.
require "json"
require "net/http"
require "uri"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
req = Net::HTTP::Post.new(uri)
req["Content-Type"] = "application/json"
req.body = JSON.generate({ "slug" => "glossary-desk" })
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode(["slug" => "glossary-desk"]));
curl_setopt($ch, CURLOPT_HTTPHEADER, ["Content-Type: application/json"]);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$guest = json_decode(curl_exec($ch), true);
curl_close($ch);
echo $guest["data"]["token"];
// Open https://glossary-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
using var http = new HttpClient();
var guestReq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/guest");
guestReq.Content = new StringContent("{\"slug\": \"glossary-desk\"}", Encoding.UTF8, "application/json");
var guestRes = await http.SendAsync(guestReq);
var guest = await guestRes.Content.ReadFromJsonAsync<JsonElement>();
Console.WriteLine(guest.GetProperty("data").GetProperty("token").GetString());
2. A tiny client
One helper that adds the headers, unwraps data and raises on error.
# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="glossary-desk"
TOKEN="$SKILLSAFE_TOKEN" # from https://glossary-desk.skillsafe.ai/tokens.html
call() { # call <path> [json-body]
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
fi
}
import json, os, urllib.error, urllib.request
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "glossary-desk"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN") # from https://glossary-desk.skillsafe.ai/tokens.html
def call(path, body=None):
"""Returns the unwrapped `data`, or raises with the API error code."""
if body is not None and not isinstance(body, dict):
raise TypeError("the request body must be a JSON object, not a bare string")
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(f"{BASE}/{path}", data=data, method="POST" if body is not None else "GET")
req.add_header("Authorization", f"Bearer {TOKEN}")
if body is not None:
req.add_header("Content-Type", "application/json")
try:
with urllib.request.urlopen(req) as r:
payload = json.load(r)
except urllib.error.HTTPError as e:
payload = json.load(e)
if not payload.get("ok"):
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code')}: {err.get('message')}")
return payload["data"]
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "glossary-desk";
const TOKEN = "aut_YOUR_TOKEN"; // from https://glossary-desk.skillsafe.ai/tokens.html
async function call(path, body) {
if (body !== undefined && (body === null || typeof body !== "object" || Array.isArray(body))) {
throw new TypeError("the request body must be a JSON object, not a bare string");
}
const res = await fetch(`${BASE}/${path}`, {
method: body ? "POST" : "GET",
headers: {
Authorization: `Bearer ${TOKEN}`,
...(body ? { "Content-Type": "application/json" } : {}),
},
body: body ? JSON.stringify(body) : undefined,
});
const payload = await res.json();
if (!payload.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
return payload.data;
}
package main
import (
"bufio"
"bytes"
"crypto/sha256"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const (
base = "https://api.skillsafe.ai/v1/app-api"
slug = "glossary-desk"
)
var token = os.Getenv("SKILLSAFE_TOKEN") // from https://glossary-desk.skillsafe.ai/tokens.html
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(path string, body any) (json.RawMessage, error) {
method := http.MethodGet
var rdr io.Reader
if body != nil {
method = http.MethodPost
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
import java.net.URI;
import java.net.http.*;
public class GlossaryDesk {
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final String SLUG = "glossary-desk";
static final String TOKEN = System.getenv().getOrDefault("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String call(String path, String jsonBody) throws Exception {
if (jsonBody != null && !jsonBody.trim().startsWith("{")) {
throw new IllegalArgumentException("the request body must be a JSON object");
}
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + TOKEN);
if (jsonBody != null) {
b.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(jsonBody));
} else {
b.GET();
}
HttpResponse<String> res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
// The envelope is always {"ok":true,"data":...} or {"ok":false,"error":...}.
return res.body();
}
}
require "json"
require "net/http"
require "uri"
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "glossary-desk"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN") # from https://glossary-desk.skillsafe.ai/tokens.html
def call(path, body = nil)
raise TypeError, "the request body must be a JSON object" if body && !body.is_a?(Hash)
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
if body
req["Content-Type"] = "application/json"
req.body = JSON.generate(body)
end
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
raise "#{payload['error']['code']}: #{payload['error']['message']}" unless payload["ok"]
payload["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "glossary-desk";
define("TOKEN", getenv("SKILLSAFE_TOKEN") ?: "aut_YOUR_TOKEN"); // from /tokens.html
function call(string $path, ?array $body = null) {
$ch = curl_init(BASE . "/" . $path);
$headers = ["Authorization: Bearer " . TOKEN];
if ($body !== null) {
$headers[] = "Content-Type: application/json";
curl_setopt($ch, CURLOPT_POST, true);
// JSON_FORCE_OBJECT is not needed here, but the body must encode as an
// object: an empty PHP array would encode as [] and be rejected.
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body, JSON_UNESCAPED_SLASHES));
}
curl_setopt($ch, CURLOPT_HTTPHEADER, $headers);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
if (empty($payload["ok"])) {
throw new RuntimeException($payload["error"]["code"] . ": " . $payload["error"]["message"]);
}
return $payload["data"];
}
using System.Net.Http.Json;
using System.Text.Json;
static class GlossaryDesk
{
const string Base = "https://api.skillsafe.ai/v1/app-api";
const string Slug = "glossary-desk";
static readonly string Token =
Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "aut_YOUR_TOKEN";
static readonly HttpClient Http = new();
public static async Task<JsonElement> Call(string path, object? body = null)
{
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post, $"{Base}/{path}");
req.Headers.Add("Authorization", $"Bearer {Token}");
if (body is not null) req.Content = JsonContent.Create(body);
var res = await Http.SendAsync(req);
var payload = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!payload.GetProperty("ok").GetBoolean())
{
var e = payload.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return payload.GetProperty("data");
}
}
3. Check the session and the balance
GET /me tells you whether the token is a guest or a person, and what the balance is.
The object is small and carries exactly three things: subject_type —
guest or user, and a guest can price a run but not start one
— subject_id, and credits, the wallet balance. There is no username in
it, so "signed in" is subject_type == "user" and nothing else. Compare
credits against min_credits from the next step before you run, so a
shortfall surfaces as your own clear message rather than a 402.
call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}
me = call("me")
signed_in = me["subject_type"] == "user"
print(me["subject_type"], me.get("credits"), "signed in" if signed_in else "guest")
const me = await call("me");
const signedIn = me.subject_type === "user";
console.log(me.subject_type, me.credits, signedIn ? "signed in" : "guest");
raw, err := call("me", nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
SubjectID string `json:"subject_id"`
Credits int `json:"credits"`
}
_ = json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits, me.SubjectType == "user")
System.out.println(call("me", null));
// {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}
// "signed in" is subject_type.equals("user") - there is no username field.
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
abort "sign in first - a guest token cannot run a lane" unless me["subject_type"] == "user"
<?php
$me = call("me");
echo $me["subject_type"], " ", $me["credits"], PHP_EOL;
$signedIn = $me["subject_type"] === "user";
var me = await GlossaryDesk.Call("me");
var signedIn = me.GetProperty("subject_type").GetString() == "user";
Console.WriteLine($"{me.GetProperty("subject_type")} {me.GetProperty("credits")} {signedIn}");
4. Price the run — free
The input object is exactly what the app's own form submits. It is always a JSON
object — never a bare string, never wrapped in an input key:
| field | type | meaning |
|---|---|---|
task | string | "glossary" or "questionnaire". The lane. Missing or unknown degrades to the closest lane by the fields present, and the reply names what it chose in lane and in notes_on_input. |
document | string, required | The pasted text: a spec, a PRD, a design note, a meeting transcript, a support thread, a policy, an RFC. This is the run's only evidence — every term, definition, ambiguity and question has to rest on words that are in it. The browser normalises line endings and clips to 60,000 characters from the middle, keeping the beginning and the end, and leaves the marker [... N characters cut from the middle of the document - the beginning and the end are kept ...] in place of the cut. Clip the same way if you send more, and keep the marker: the prompt keys on it, refuses to claim anything about the missing middle, and mentions the cut in notes_on_input. |
context_hint | string, optional | One or two sentences about what the project is and who wrote the document — "A B2B SaaS product. The PRD was written by a PM and edited by two engineers." The browser caps it at 2,000 characters. |
decisions | string, optional | Settled answers from an earlier questionnaire, as free text. Authoritative: a term a decision settles comes back firm with the decision as its evidence, an ambiguity the decisions resolve is not raised again, and a decision that contradicts the document still wins — with the contradiction noted in notes_on_input. Capped at 8,000 characters. |
prescan_facts | object | What a local scanner found, for free, before the run. Shape below. Empty arrays are legitimate — read the honesty note. The platform's input check declares scalar fields only, so /estimate answers an object-valued prescan_facts or needs with an advisory unknown field warning; that warning is expected and the field still reaches the model. Only task and document are required. |
retry_note | string, optional | Send only on a retry, when a previous reply failed to parse or came back truncated. The instruction is obeyed exactly. The web app adds it automatically on its one automatic retry. |
| questionnaire lane only | ||
recipient | string | One named person, their role, and how you relate to them — "Maya, Head of Billing. She owns the pricing and refund rules and signs off invoicing changes. I am the PM on the plan-change feature." Without it the questionnaire is pitched at an unnamed colleague, which is weaker. Capped at 2,000 characters. |
needs | object[] | [{"id": "N-001", "text": "…"}] — what you need back, one decision per entry, ids sequential from N-001. The browser numbers up to 20 lines of typed text (6,000 characters) and strips any bullet or number the user typed. Every id you send comes back exactly once in coverage. Omitting needs is allowed: the model derives them from the document's own open items and says so, and coverage comes back empty. |
deadline | string, optional | When you need the answers and how long it should take — "Answers by Thursday; 20 minutes should be enough." Capped at 2,000 characters. |
prescan_facts, honestly
In the browser this object is computed for free by a local scanner before the run:
{
"stats": { "words": 210, "sentences": 14, "headings": 5, "paragraphs": 9 },
"candidates": [
{ "id": "C-001", "term": "customer", "count": 4,
"forms": ["customer", "customers"], "first_seen": "line 5", "strong": true }
],
"clusters": [
{ "id": "K-001", "terms": ["customer", "user", "client", "account"],
"reason": "known synonym family, all four appear" }
],
"redefinitions": [
{ "id": "R-001", "term": "cancellation", "kind": "conflicting_rules",
"quotes": ["Cancellation takes effect immediately", "keeps read-only access until the end of the period"] }
],
"clipped": { "cut": 0 }
}
candidates are the repeated noun phrases the scanner thinks might be terms, with their
occurrence count and surface forms; clusters are groups of words that look like
competing names for one thing; redefinitions are places the same word is defined twice
or where two rules about it conflict — kind is defined_twice or
conflicting_rules. clipped.cut is how many characters the middle clip
removed.
An API caller does not have to reproduce any of that. Sending
{"stats": {...}, "candidates": [], "clusters": [], "redefinitions": [], "clipped": {"cut": 0}}
is legitimate and the run still works — the model reads document either way. Be
honest with yourself about what you give up: the glossary lane's coverage list
comes back empty, because it is one entry per candidate id, and the client-side
reconciliation — which checks that every candidate came back with a status, that every cluster and
redefinition was addressed somewhere, and that no id you did not send appears — then has nothing to
check. You lose the accountability, not the glossary. The questionnaire lane is unaffected: its
coverage is keyed on needs, not on candidates, so empty candidate arrays
cost it nothing.
What makes the facts worth sending is that contract: every id in
prescan_facts.candidates comes back exactly once in coverage,
with a status of defined, avoid, excluded,
ambiguity or not_a_term. A mechanical scanner is allowed to be wrong, and
not_a_term is the honest answer for a phrase that only looked like a term — that is a
different thing from silence. A candidate that never appears at all is a failed run, not a passing
one.
/estimate creates no job and charges nothing. It returns the model
binding — model is gpt-5.6-terra, model_alias is
gpt-terra, markup_bps is 1000 — and the reservation:
hold_credits is what gets held, and min_credits is the balance you must
clear to start at all. The hold is a reservation, not the price. It prices the full output
cap, so the charged_credits you see on the settled job is usually far lower. Budget
against hold_credits, report against charged_credits.
Price each lane separately. The questionnaire body is shorter than a glossary body, and the same
document with an empty prescan_facts prices differently from one carrying forty
candidates — the facts are input tokens like everything else.
# Lane A - the glossary. One document, the facts a local scan found, no decisions yet.
INPUT='{"task": "glossary", "document": "# Subscription billing - PRD v0.3\n\n## Plans\nWe sell three tiers: Starter, Team and Business. A user picks a package at sign-up and can change tier at any time from the billing page.\n\n## Cancellation\nA client can cancel from the billing page. Cancellation takes effect immediately: access ends and no further invoices are raised.\n\n## Invoices\nEvery charge produces an invoice, emailed to the billing contact. Failed payments retry three times over seven days, after which the subscription is cancelled - the customer keeps read-only access until the end of the period they already paid for.\n\n## Seats\nTeam and Business are priced per seat. A seat is an invited member who has accepted; pending invitations do not count.", "context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.", "decisions": "", "prescan_facts": {"stats": {"words": 210, "sentences": 14, "headings": 5, "paragraphs": 9}, "candidates": [{"id": "C-001", "term": "customer", "count": 4, "forms": ["customer", "customers"], "first_seen": "line 14", "strong": true}, {"id": "C-002", "term": "seat", "count": 3, "forms": ["seat", "seats"], "first_seen": "line 17", "strong": true}, {"id": "C-003", "term": "billing page", "count": 3, "forms": ["billing page"], "first_seen": "line 5", "strong": false}], "clusters": [{"id": "K-001", "terms": ["customer", "user", "client"], "reason": "known synonym family, all three appear"}], "redefinitions": [{"id": "R-001", "term": "cancellation", "kind": "conflicting_rules", "quotes": ["Cancellation takes effect immediately: access ends", "the customer keeps read-only access until the end of the period they already paid for"]}], "clipped": {"cut": 0}}}'
call estimate "$INPUT"
# {"ok":true,"data":{"model":"gpt-5.6-terra","model_alias":"gpt-terra",
# "markup_bps":1000,"hold_credits":2652,"min_credits":310}}
#
# estimate is FREE. It creates no job and charges nothing. hold_credits is what
# gets RESERVED; charged_credits on the settled job is normally much lower.
# Lane B - the questionnaire over the same document, aimed at one person, with
# two things only she can settle. Empty prescan arrays are fine here: this lane's
# coverage is keyed on `needs`, not on candidates.
Q_INPUT='{"task": "questionnaire", "document": "# Subscription billing - PRD v0.3\n\n## Cancellation\nA client can cancel from the billing page. Cancellation takes effect immediately: access ends and no further invoices are raised.\n\n## Invoices\nFailed payments retry three times over seven days, after which the subscription is cancelled - the customer keeps read-only access until the end of the period they already paid for.", "context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.", "decisions": "", "recipient": "Maya, Head of Billing. She owns the pricing and refund rules and signs off invoicing changes. I am the PM on the plan-change feature.", "needs": [{"id": "N-001", "text": "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both."}, {"id": "N-002", "text": "Whether customer, user and client are one thing or several, and which word Billing uses on invoices."}], "deadline": "Answers by Thursday; 20 minutes should be enough.", "prescan_facts": {"stats": {"words": 95, "sentences": 5, "headings": 3, "paragraphs": 4}, "candidates": [], "clusters": [], "redefinitions": [], "clipped": {"cut": 0}}}'
call estimate "$Q_INPUT"
DOC = """# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
## Seats
Team and Business are priced per seat. A seat is an invited member who has accepted;
pending invitations do not count."""
HINT = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers."
FACTS = {
"stats": {"words": 210, "sentences": 14, "headings": 5, "paragraphs": 9},
"candidates": [
{"id": "C-001", "term": "customer", "count": 4,
"forms": ["customer", "customers"], "first_seen": "line 14", "strong": True},
{"id": "C-002", "term": "seat", "count": 3,
"forms": ["seat", "seats"], "first_seen": "line 17", "strong": True},
{"id": "C-003", "term": "billing page", "count": 3,
"forms": ["billing page"], "first_seen": "line 5", "strong": False},
],
"clusters": [
{"id": "K-001", "terms": ["customer", "user", "client"],
"reason": "known synonym family, all three appear"},
],
"redefinitions": [
{"id": "R-001", "term": "cancellation", "kind": "conflicting_rules",
"quotes": ["Cancellation takes effect immediately: access ends",
"keeps read-only access until the end of the period they already paid for"]},
],
"clipped": {"cut": 0},
}
# Lane A - the glossary.
INPUT = {
"task": "glossary",
"document": DOC,
"context_hint": HINT,
"decisions": "", # answers collected earlier go here, and outrank the document
"prescan_facts": FACTS,
}
# Lane B - the questionnaire over the same document. Empty candidate arrays are
# legitimate: this lane's coverage is keyed on `needs`.
Q_INPUT = {
"task": "questionnaire",
"document": DOC,
"context_hint": HINT,
"decisions": "",
"recipient": ("Maya, Head of Billing. She owns the pricing and refund rules and signs "
"off invoicing changes. I am the PM on the plan-change feature."),
"needs": [
{"id": "N-001", "text": "Whether a cancellation ends access immediately or at the "
"end of the paid period - the PRD says both."},
{"id": "N-002", "text": "Whether customer, user and client are one thing or several, "
"and which word Billing uses on invoices."},
],
"deadline": "Answers by Thursday; 20 minutes should be enough.",
"prescan_facts": {"stats": FACTS["stats"], "candidates": [], "clusters": [],
"redefinitions": [], "clipped": {"cut": 0}},
}
est = call("estimate", INPUT)
print(est["model"], est["model_alias"], est["markup_bps"]) # gpt-5.6-terra gpt-terra 1000
print(est["hold_credits"], est["min_credits"])
# estimate is free: no job is created and nothing is charged. The hold is a
# reservation against the full output cap, not the price of the run.
print(call("estimate", Q_INPUT)["hold_credits"]) # price each lane separately
const DOC = `# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
## Seats
Team and Business are priced per seat. A seat is an invited member who has accepted;
pending invitations do not count.`;
const HINT = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.";
const FACTS = {
stats: { words: 210, sentences: 14, headings: 5, paragraphs: 9 },
candidates: [
{ id: "C-001", term: "customer", count: 4, forms: ["customer", "customers"], first_seen: "line 14", strong: true },
{ id: "C-002", term: "seat", count: 3, forms: ["seat", "seats"], first_seen: "line 17", strong: true },
{ id: "C-003", term: "billing page", count: 3, forms: ["billing page"], first_seen: "line 5", strong: false },
],
clusters: [
{ id: "K-001", terms: ["customer", "user", "client"], reason: "known synonym family, all three appear" },
],
redefinitions: [
{ id: "R-001", term: "cancellation", kind: "conflicting_rules",
quotes: ["Cancellation takes effect immediately: access ends",
"keeps read-only access until the end of the period they already paid for"] },
],
clipped: { cut: 0 },
};
// Lane A - the glossary.
const INPUT = {
task: "glossary",
document: DOC,
context_hint: HINT,
decisions: "", // answers collected earlier go here, and outrank the document
prescan_facts: FACTS,
};
// Lane B - the questionnaire over the same document.
const Q_INPUT = {
task: "questionnaire",
document: DOC,
context_hint: HINT,
decisions: "",
recipient:
"Maya, Head of Billing. She owns the pricing and refund rules and signs off " +
"invoicing changes. I am the PM on the plan-change feature.",
needs: [
{ id: "N-001", text: "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both." },
{ id: "N-002", text: "Whether customer, user and client are one thing or several, and which word Billing uses on invoices." },
],
deadline: "Answers by Thursday; 20 minutes should be enough.",
// Empty candidate arrays are legitimate; this lane's coverage is keyed on `needs`.
prescan_facts: { stats: FACTS.stats, candidates: [], clusters: [], redefinitions: [], clipped: { cut: 0 } },
};
const est = await call("estimate", INPUT);
console.log(est.model, est.model_alias, est.markup_bps); // gpt-5.6-terra gpt-terra 1000
console.log(est.hold_credits, est.min_credits);
// estimate is free: no job is created and nothing is charged. hold_credits is a
// reservation against the output cap; charged_credits is normally far lower.
console.log((await call("estimate", Q_INPUT)).hold_credits);
const doc = `# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
## Seats
Team and Business are priced per seat. A seat is an invited member who has accepted;
pending invitations do not count.`
const hint = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers."
stats := map[string]any{"words": 210, "sentences": 14, "headings": 5, "paragraphs": 9}
facts := map[string]any{
"stats": stats,
"candidates": []any{
map[string]any{"id": "C-001", "term": "customer", "count": 4,
"forms": []string{"customer", "customers"}, "first_seen": "line 14", "strong": true},
map[string]any{"id": "C-002", "term": "seat", "count": 3,
"forms": []string{"seat", "seats"}, "first_seen": "line 17", "strong": true},
},
"clusters": []any{
map[string]any{"id": "K-001", "terms": []string{"customer", "user", "client"},
"reason": "known synonym family, all three appear"},
},
"redefinitions": []any{
map[string]any{"id": "R-001", "term": "cancellation", "kind": "conflicting_rules",
"quotes": []string{"Cancellation takes effect immediately: access ends",
"keeps read-only access until the end of the period they already paid for"}},
},
"clipped": map[string]any{"cut": 0},
}
// Lane A - the glossary.
input := map[string]any{
"task": "glossary",
"document": doc,
"context_hint": hint,
"decisions": "", // answers collected earlier go here, and outrank the document
"prescan_facts": facts,
}
// Lane B - the questionnaire over the same document. Empty candidate arrays are
// legitimate: this lane's coverage is keyed on needs.
qInput := map[string]any{
"task": "questionnaire",
"document": doc,
"context_hint": hint,
"recipient": "Maya, Head of Billing. She owns the pricing and refund rules and signs off " +
"invoicing changes. I am the PM on the plan-change feature.",
"needs": []any{
map[string]string{"id": "N-001", "text": "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both."},
map[string]string{"id": "N-002", "text": "Whether customer, user and client are one thing or several, and which word Billing uses on invoices."},
},
"deadline": "Answers by Thursday; 20 minutes should be enough.",
"prescan_facts": map[string]any{"stats": stats, "candidates": []any{},
"clusters": []any{}, "redefinitions": []any{}, "clipped": map[string]any{"cut": 0}},
}
raw, err := call("estimate", input)
if err != nil {
panic(err)
}
fmt.Println(string(raw)) // free: no job, no charge; hold_credits is a reservation
if q, err := call("estimate", qInput); err == nil {
fmt.Println(string(q)) // price each lane separately
}
// estimate is free: no job is created and nothing is charged. The data object
// carries model (gpt-5.6-terra), model_alias (gpt-terra), markup_bps (1000),
// hold_credits and min_credits. The hold is a reservation against the full
// output cap, so the settled charge is normally far lower.
String doc = """
# Subscription billing - PRD v0.3
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Failed payments retry three times over seven days, after which the subscription is
cancelled - the customer keeps read-only access until the end of the period they
already paid for.""".replace("\n", "\\n").replace("\"", "\\\"");
// Lane A - the glossary.
String input = """
{
"task": "glossary",
"document": "%s",
"context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.",
"decisions": "",
"prescan_facts": {
"stats": { "words": 210, "sentences": 14, "headings": 5, "paragraphs": 9 },
"candidates": [
{ "id": "C-001", "term": "customer", "count": 4,
"forms": ["customer", "customers"], "first_seen": "line 14", "strong": true },
{ "id": "C-002", "term": "seat", "count": 3,
"forms": ["seat", "seats"], "first_seen": "line 17", "strong": true }
],
"clusters": [
{ "id": "K-001", "terms": ["customer", "user", "client"],
"reason": "known synonym family, all three appear" }
],
"redefinitions": [
{ "id": "R-001", "term": "cancellation", "kind": "conflicting_rules",
"quotes": ["Cancellation takes effect immediately: access ends"] }
],
"clipped": { "cut": 0 }
}
}""".formatted(doc);
// Lane B - the questionnaire over the same document, aimed at one person.
String qInput = """
{
"task": "questionnaire",
"document": "%s",
"context_hint": "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.",
"recipient": "Maya, Head of Billing. She owns the pricing and refund rules and signs off invoicing changes. I am the PM on the plan-change feature.",
"needs": [
{ "id": "N-001", "text": "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both." },
{ "id": "N-002", "text": "Whether customer, user and client are one thing or several, and which word Billing uses on invoices." }
],
"deadline": "Answers by Thursday; 20 minutes should be enough.",
"prescan_facts": { "stats": { "words": 95, "sentences": 5, "headings": 3, "paragraphs": 4 },
"candidates": [], "clusters": [], "redefinitions": [], "clipped": { "cut": 0 } }
}""".formatted(doc);
System.out.println(call("estimate", input));
System.out.println(call("estimate", qInput));
DOC = <<~TEXT
# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
TEXT
HINT = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers."
STATS = { "words" => 210, "sentences" => 14, "headings" => 5, "paragraphs" => 9 }
facts = {
"stats" => STATS,
"candidates" => [
{ "id" => "C-001", "term" => "customer", "count" => 4,
"forms" => ["customer", "customers"], "first_seen" => "line 14", "strong" => true },
{ "id" => "C-002", "term" => "seat", "count" => 3,
"forms" => ["seat", "seats"], "first_seen" => "line 17", "strong" => true }
],
"clusters" => [
{ "id" => "K-001", "terms" => ["customer", "user", "client"],
"reason" => "known synonym family, all three appear" }
],
"redefinitions" => [
{ "id" => "R-001", "term" => "cancellation", "kind" => "conflicting_rules",
"quotes" => ["Cancellation takes effect immediately: access ends"] }
],
"clipped" => { "cut" => 0 }
}
# Lane A - the glossary.
input = { "task" => "glossary", "document" => DOC, "context_hint" => HINT,
"decisions" => "", "prescan_facts" => facts }
# Lane B - the questionnaire. Empty candidate arrays are legitimate here.
q_input = {
"task" => "questionnaire", "document" => DOC, "context_hint" => HINT,
"recipient" => "Maya, Head of Billing. She owns the pricing and refund rules and " \
"signs off invoicing changes. I am the PM on the plan-change feature.",
"needs" => [
{ "id" => "N-001", "text" => "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both." },
{ "id" => "N-002", "text" => "Whether customer, user and client are one thing or several, and which word Billing uses on invoices." }
],
"deadline" => "Answers by Thursday; 20 minutes should be enough.",
"prescan_facts" => { "stats" => STATS, "candidates" => [], "clusters" => [],
"redefinitions" => [], "clipped" => { "cut" => 0 } }
}
est = call("estimate", input)
puts "#{est['model']} #{est['model_alias']} hold=#{est['hold_credits']} min=#{est['min_credits']}"
puts call("estimate", q_input)["hold_credits"]
# estimate is free: no job is created and nothing is charged.
<?php
$doc = <<<TEXT
# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
TEXT;
$hint = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.";
$stats = ["words" => 210, "sentences" => 14, "headings" => 5, "paragraphs" => 9];
$facts = [
"stats" => $stats,
"candidates" => [
["id" => "C-001", "term" => "customer", "count" => 4,
"forms" => ["customer", "customers"], "first_seen" => "line 14", "strong" => true],
["id" => "C-002", "term" => "seat", "count" => 3,
"forms" => ["seat", "seats"], "first_seen" => "line 17", "strong" => true],
],
"clusters" => [
["id" => "K-001", "terms" => ["customer", "user", "client"],
"reason" => "known synonym family, all three appear"],
],
"redefinitions" => [
["id" => "R-001", "term" => "cancellation", "kind" => "conflicting_rules",
"quotes" => ["Cancellation takes effect immediately: access ends"]],
],
"clipped" => ["cut" => 0],
];
// Lane A - the glossary.
$input = ["task" => "glossary", "document" => $doc, "context_hint" => $hint,
"decisions" => "", "prescan_facts" => $facts];
// Lane B - the questionnaire. Empty candidate arrays are legitimate here.
$qInput = [
"task" => "questionnaire", "document" => $doc, "context_hint" => $hint,
"recipient" => "Maya, Head of Billing. She owns the pricing and refund rules and signs "
. "off invoicing changes. I am the PM on the plan-change feature.",
"needs" => [
["id" => "N-001", "text" => "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both."],
["id" => "N-002", "text" => "Whether customer, user and client are one thing or several, and which word Billing uses on invoices."],
],
"deadline" => "Answers by Thursday; 20 minutes should be enough.",
"prescan_facts" => ["stats" => $stats, "candidates" => [], "clusters" => [],
"redefinitions" => [], "clipped" => ["cut" => 0]],
];
$est = call("estimate", $input);
echo $est["model"], " ", $est["hold_credits"], " ", $est["min_credits"], PHP_EOL;
echo call("estimate", $qInput)["hold_credits"], PHP_EOL;
// estimate is free: no job is created and nothing is charged.
var doc = """
# Subscription billing - PRD v0.3
## Plans
We sell three tiers: Starter, Team and Business. A user picks a package at sign-up and
can change tier at any time from the billing page.
## Cancellation
A client can cancel from the billing page. Cancellation takes effect immediately:
access ends and no further invoices are raised.
## Invoices
Every charge produces an invoice, emailed to the billing contact. Failed payments retry
three times over seven days, after which the subscription is cancelled - the customer
keeps read-only access until the end of the period they already paid for.
""";
var hint = "A B2B SaaS product. The PRD was written by a PM and edited by two engineers.";
var stats = new { words = 210, sentences = 14, headings = 5, paragraphs = 9 };
// Lane A - the glossary.
var input = new
{
task = "glossary",
document = doc,
context_hint = hint,
decisions = "",
prescan_facts = new
{
stats,
candidates = new[]
{
new { id = "C-001", term = "customer", count = 4,
forms = new[] { "customer", "customers" }, first_seen = "line 14", strong = true },
new { id = "C-002", term = "seat", count = 3,
forms = new[] { "seat", "seats" }, first_seen = "line 17", strong = true },
},
clusters = new[]
{
new { id = "K-001", terms = new[] { "customer", "user", "client" },
reason = "known synonym family, all three appear" },
},
redefinitions = new[]
{
new { id = "R-001", term = "cancellation", kind = "conflicting_rules",
quotes = new[] { "Cancellation takes effect immediately: access ends" } },
},
clipped = new { cut = 0 },
},
};
// Lane B - the questionnaire over the same document.
var qInput = new
{
task = "questionnaire",
document = doc,
context_hint = hint,
recipient = "Maya, Head of Billing. She owns the pricing and refund rules and signs off " +
"invoicing changes. I am the PM on the plan-change feature.",
needs = new[]
{
new { id = "N-001", text = "Whether a cancellation ends access immediately or at the end of the paid period - the PRD says both." },
new { id = "N-002", text = "Whether customer, user and client are one thing or several, and which word Billing uses on invoices." },
},
deadline = "Answers by Thursday; 20 minutes should be enough.",
prescan_facts = new { stats, candidates = Array.Empty<object>(), clusters = Array.Empty<object>(),
redefinitions = Array.Empty<object>(), clipped = new { cut = 0 } },
};
var est = await GlossaryDesk.Call("estimate", input);
Console.WriteLine(est.GetProperty("model").GetString()); // gpt-5.6-terra
Console.WriteLine(est.GetProperty("hold_credits").GetInt32()); // a reservation, not the price
Console.WriteLine((await GlossaryDesk.Call("estimate", qInput)).GetProperty("hold_credits").GetInt32());
// estimate is free: no job is created and nothing is charged.
5. Run it, then poll
POST /run returns a job_id; poll GET jobs/{job_id} until
status is succeeded or failed. The result JSON is the string
at data.output.output. The terminal job also carries charged_credits —
the real price — and the truncated flag.
Always send an Idempotency-Key. The web app builds it as
glossary-desk:<task>:<input hash>:a<attempt> and so should you. Three
parts, three reasons:
- the task, because two lanes over one document are two different runs and must never collide on one key;
- a hash of the input — of what a person actually typed:
task,document,context_hint,decisions,recipient, theneedstext anddeadline. The hash deliberately excludesprescan_facts, so re-running the same paste with a slightly better local scan is still the same run; - an attempt counter, because a retry with a changed body — a
retry_noteadded after a reply failed to parse — must not replay the old key. Replaying a key with a different body is a 409conflict.
A retried request carrying the same key returns the same job instead of billing a second run, which is what makes a CI retry safe after a network blip.
# The key is slug:task:hash:attempt. A retried request with the same key returns
# the SAME job instead of billing a second run.
KEY="glossary-desk:glossary:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
# Poll until the job reaches a terminal status.
while :; do
OUT=$(call "jobs/$JOB")
STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
sleep 2
done
# The terminal job looks like this:
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
# "output":{"output":"{\"lane\":\"glossary\",\"title\":\"Subscription billing\", ...}"},
# "charged_credits":588,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])'
import hashlib, time
def idem_key(inp, attempt=1):
"""slug:task:hash:attempt - the hash covers what a person typed, not the scan."""
typed = [inp.get("task", ""), inp.get("document", ""), inp.get("context_hint", ""),
inp.get("decisions", ""), inp.get("recipient", ""),
"\n".join(n["text"] for n in inp.get("needs", [])), inp.get("deadline", "")]
digest = hashlib.sha256("␟".join(s.strip() for s in typed).encode()).hexdigest()[:16]
return f"glossary-desk:{inp.get('task', 'glossary')}:{digest}:a{attempt}"
key = idem_key(INPUT)
req = urllib.request.Request(f"{BASE}/run", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
with urllib.request.urlopen(req) as r:
job_id = json.load(r)["data"]["job_id"]
while True:
job = call(f"jobs/{job_id}")
if job["status"] == "succeeded":
break
if job["status"] == "failed":
raise RuntimeError(job.get("error"))
time.sleep(2)
result = json.loads(job["output"]["output"])
print(result["lane"], result["title"], "-", result.get("verdict"))
print(len(result.get("terms", [])), "terms,", len(result.get("ambiguities", [])), "ambiguities")
print("charged", job.get("charged_credits"), "truncated", job.get("truncated"))
import { createHash } from "node:crypto";
// slug:task:hash:attempt - the hash covers what a person typed, not the scan.
function idemKey(inp, attempt = 1) {
const typed = [inp.task ?? "", inp.document ?? "", inp.context_hint ?? "", inp.decisions ?? "",
inp.recipient ?? "", (inp.needs ?? []).map((n) => n.text).join("\n"),
inp.deadline ?? ""];
const digest = createHash("sha256").update(typed.map((s) => s.trim()).join("␟"))
.digest("hex").slice(0, 16);
return `glossary-desk:${inp.task ?? "glossary"}:${digest}:a${attempt}`;
}
const key = idemKey(INPUT);
const started = await fetch(`${BASE}/run`, {
method: "POST",
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify(INPUT),
}).then((r) => r.json());
let job = started.data;
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`jobs/${job.job_id}`);
}
if (job.status === "failed") throw new Error(JSON.stringify(job.error));
const result = JSON.parse(job.output.output);
console.log(result.lane, result.title, result.verdict, job.charged_credits);
// slug:task:hash:attempt. A retried request with the same key returns the SAME
// job instead of billing a second run.
body, _ := json.Marshal(input)
sum := sha256.Sum256(body)
key := fmt.Sprintf("glossary-desk:glossary:%x:a1", sum[:8])
req, _ := http.NewRequest(http.MethodPost, base+"/run", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, _ := http.DefaultClient.Do(req)
defer res.Body.Close()
var started struct {
Data struct {
JobID string `json:"job_id"`
} `json:"data"`
}
_ = json.NewDecoder(res.Body).Decode(&started)
for {
raw, err := call("jobs/"+started.Data.JobID, nil)
if err != nil {
panic(err)
}
var job struct {
Status string `json:"status"`
Output struct {
Output string `json:"output"`
} `json:"output"`
ChargedCredits int `json:"charged_credits"`
Truncated bool `json:"truncated"`
}
_ = json.Unmarshal(raw, &job)
if job.Status == "succeeded" {
fmt.Println(job.Output.Output) // the glossary JSON, as a string
fmt.Println(job.ChargedCredits, job.Truncated)
break
}
if job.Status == "failed" {
panic("run failed")
}
time.Sleep(2 * time.Second)
}
// slug:task:hash:attempt. A retried request with the same key returns the SAME
// job instead of billing a second run; a CHANGED body under an old key is a 409.
var digest = java.security.MessageDigest.getInstance("SHA-256")
.digest(input.getBytes(java.nio.charset.StandardCharsets.UTF_8));
var key = "glossary-desk:glossary:"
+ java.util.HexFormat.of().formatHex(digest).substring(0, 16) + ":a1";
var start = HttpRequest.newBuilder(URI.create(BASE + "/run"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.POST(HttpRequest.BodyPublishers.ofString(input))
.build();
String started = HTTP.send(start, HttpResponse.BodyHandlers.ofString()).body();
// Parse job_id out of `started`, then poll GET jobs/{job_id} every two seconds
// until status is "succeeded" or "failed". The result JSON is data.output.output,
// and the terminal job also carries charged_credits and truncated.
System.out.println(started);
require "digest"
# slug:task:hash:attempt - the hash covers what a person typed, not the scan.
def idem_key(inp, attempt = 1)
typed = [inp["task"], inp["document"], inp["context_hint"], inp["decisions"],
inp["recipient"], (inp["needs"] || []).map { |n| n["text"] }.join("\n"),
inp["deadline"]].map { |v| v.to_s.strip }
"glossary-desk:#{inp['task']}:#{Digest::SHA256.hexdigest(typed.join("␟"))[0, 16]}:a#{attempt}"
end
key = idem_key(input)
uri = URI("#{BASE}/run")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = JSON.generate(input)
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
job_id = JSON.parse(res.body)["data"]["job_id"]
loop do
job = call("jobs/#{job_id}")
if job["status"] == "succeeded"
result = JSON.parse(job["output"]["output"])
puts "#{result['lane']} #{result['title']} #{result['verdict']}"
puts "charged=#{job['charged_credits']} truncated=#{job['truncated']}"
break
end
raise "run failed" if job["status"] == "failed"
sleep 2
end
<?php
// slug:task:hash:attempt - the hash covers what a person typed, not the scan.
function idemKey(array $inp, int $attempt = 1): string {
$typed = [$inp["task"] ?? "", $inp["document"] ?? "", $inp["context_hint"] ?? "",
$inp["decisions"] ?? "", $inp["recipient"] ?? "",
implode("\n", array_column($inp["needs"] ?? [], "text")), $inp["deadline"] ?? ""];
$digest = substr(hash("sha256", implode("\u{241f}", array_map("trim", $typed))), 0, 16);
return "glossary-desk:" . ($inp["task"] ?? "glossary") . ":{$digest}:a{$attempt}";
}
$key = idemKey($input);
$ch = curl_init(BASE . "/run");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($input));
curl_setopt($ch, CURLOPT_HTTPHEADER, [
"Authorization: Bearer " . TOKEN,
"Content-Type: application/json",
"Idempotency-Key: " . $key,
]);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$jobId = json_decode(curl_exec($ch), true)["data"]["job_id"];
curl_close($ch);
while (true) {
$job = call("jobs/" . $jobId);
if ($job["status"] === "succeeded") {
$result = json_decode($job["output"]["output"], true);
echo $result["lane"], " ", $result["title"], " ", $result["verdict"] ?? "", PHP_EOL;
echo "charged=", $job["charged_credits"], " truncated=", var_export($job["truncated"], true), PHP_EOL;
break;
}
if ($job["status"] === "failed") { throw new RuntimeException("run failed"); }
sleep(2);
}
using System.Security.Cryptography;
using System.Text;
// slug:task:hash:attempt. A retried request with the same key returns the SAME
// job instead of billing a second run.
var json = JsonSerializer.Serialize(input);
var digest = Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(json)))[..16].ToLowerInvariant();
var key = $"glossary-desk:glossary:{digest}:a1";
var run = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run");
run.Headers.Add("Authorization", "Bearer aut_YOUR_TOKEN");
run.Headers.Add("Idempotency-Key", key);
run.Content = JsonContent.Create(input);
// POST it, read data.job_id, then poll GET jobs/{job_id} every two seconds until
// status is "succeeded" or "failed". The result JSON is data.output.output, and
// the terminal job also carries charged_credits and truncated.
6. Or stream it
POST /run-stream is the same call over server-sent events, and it takes the same
Idempotency-Key. Each delta event carries {"text": "..."}, a
chunk of the result JSON, and the final done event carries status,
charged_credits — the real price, normally a fraction of the hold — and the
truncated flag.
Read the SSE yourself. What a client receives depends on where it is: a
command-line reader like the ones below gets real delta events, while the same
endpoint sends a page in a browser tick heartbeats instead — so a JavaScript callback
wired to deltas never fires there, and any progress display, streaming preview or partial-recovery
path built on it is dead code in a browser. Parse the event stream in your own reader, as the
samples here do, and treat a run with no deltas at all as normal rather than as a stall: wait for
done, or fall back to /run and polling.
The practical tip: the web app does not parse the partial JSON to drive its progress display, it
watches for key names arriving in the accumulating text. In the glossary lane the appearance of
"terms", then "ambiguities", then "coverage", then
"context_md" is what advances the stage from choosing the canonical terms to quoting
the evidence, reconciling the scan, and finally writing the CONTEXT.md. In the
questionnaire lane the sequence is "purpose", "context_paragraph",
"themes", "coverage", "questionnaire_md". Substring matching
on the quoted key name is enough, and it costs nothing.
# Server-sent events. From the command line each `delta` carries a chunk of the
# JSON; the final `done` event carries the status, charged_credits and truncated.
curl -N -X POST "$BASE/run-stream" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-H "Accept: text/event-stream" \
-d "$INPUT"
# event: job {"job_id":"job_..."}
# event: delta {"text":"{\"lane\":\"glossary\",\"title\":\"Subscription"}
# event: delta {"text":" billing - PRD v0.3\",\"verdict\":\"contradictory\","}
# event: done {"status":"succeeded","charged_credits":588,"truncated":false}
#
# A page in a browser gets `tick` heartbeats here instead of `delta` events, so
# read the stream in your own reader rather than relying on a delta callback.
# Server-sent events: the result arrives in chunks, so a UI can show progress.
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
req.add_header("Accept", "text/event-stream")
raw = ""
done = {}
event = None
stage = "reading the document"
with urllib.request.urlopen(req) as stream:
for line in stream:
line = line.decode().rstrip("\n")
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: ") and event == "delta":
raw += json.loads(line[6:]).get("text", "")
# The arrival of a key name is the progress signal the web app uses.
if '"context_md"' in raw:
stage = "writing CONTEXT.md"
elif '"coverage"' in raw:
stage = "reconciling the scan and setting the verdict"
elif '"ambiguities"' in raw:
stage = "finding the ambiguities and quoting the evidence"
elif '"terms"' in raw:
stage = "choosing the canonical terms"
elif line.startswith("data: ") and event == "done":
done = json.loads(line[6:])
elif line.startswith("event: tick"):
pass # heartbeat, not content - browsers get these instead of deltas
result = json.loads(raw[raw.index("{"):raw.rindex("}") + 1])
print(stage, result["verdict"], len(result["terms"]), "terms", done.get("charged_credits"))
// Server-sent events, read by hand. Note that in a BROWSER this endpoint sends
// `tick` heartbeats rather than `delta` events, so the accumulating-text trick
// below only works from a script (node, deno, a worker you control).
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
Accept: "text/event-stream",
},
body: JSON.stringify(INPUT),
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
let raw = "";
let done = {};
let event = null;
let stage = "reading the document";
while (true) {
const chunk = await reader.read();
if (chunk.done) break;
buffer += decoder.decode(chunk.value, { stream: true });
const lines = buffer.split("\n");
buffer = lines.pop();
for (const line of lines) {
if (line.startsWith("event: ")) event = line.slice(7);
else if (line.startsWith("data: ") && event === "delta") {
raw += JSON.parse(line.slice(6)).text ?? "";
if (raw.includes('"context_md"')) stage = "writing CONTEXT.md";
else if (raw.includes('"coverage"')) stage = "reconciling the scan";
else if (raw.includes('"ambiguities"')) stage = "finding the ambiguities";
else if (raw.includes('"terms"')) stage = "choosing the canonical terms";
} else if (line.startsWith("data: ") && event === "done") {
done = JSON.parse(line.slice(6));
}
}
}
const result = JSON.parse(raw.slice(raw.indexOf("{"), raw.lastIndexOf("}") + 1));
console.log(stage, result.verdict, result.terms.length, "terms", done.charged_credits);
// Server-sent events: the result arrives in chunks, so a UI can show progress.
req, _ = http.NewRequest(http.MethodPost, base+"/run-stream", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
req.Header.Set("Accept", "text/event-stream")
res, _ = http.DefaultClient.Do(req)
defer res.Body.Close()
var raw strings.Builder
var event string
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 0, 64*1024), 4*1024*1024)
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = strings.TrimPrefix(line, "event: ")
case strings.HasPrefix(line, "data: ") && event == "delta":
var d struct {
Text string `json:"text"`
}
_ = json.Unmarshal([]byte(strings.TrimPrefix(line, "data: ")), &d)
raw.WriteString(d.Text)
// "terms", "ambiguities", "coverage", "context_md" advance the stage.
case strings.HasPrefix(line, "data: ") && event == "done":
fmt.Println(strings.TrimPrefix(line, "data: ")) // status, charged_credits, truncated
case event == "tick":
// heartbeat only - a browser gets these INSTEAD of deltas
}
}
fmt.Println(raw.String())
// Server-sent events: the result arrives in chunks, so a UI can show progress.
var stream = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.header("Accept", "text/event-stream")
.POST(HttpRequest.BodyPublishers.ofString(input))
.build();
StringBuilder raw = new StringBuilder();
String[] event = { null };
HTTP.send(stream, HttpResponse.BodyHandlers.ofLines()).body().forEach(line -> {
if (line.startsWith("event: ")) event[0] = line.substring(7);
else if (line.startsWith("data: ") && "delta".equals(event[0])) {
raw.append(line.substring(6)); // each data line is {"text":"..."} - decode and append .text
}
});
System.out.println(raw);
// Watch the accumulating text for "terms", "ambiguities", "coverage" and
// "context_md" to advance a progress display. The final `done` event carries
// status, charged_credits and truncated. A browser client sees `tick`
// heartbeats here instead of deltas, so read the stream yourself.
# Server-sent events: the result arrives in chunks, so a UI can show progress.
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req["Accept"] = "text/event-stream"
req.body = JSON.generate(input)
raw = +""
event = nil
Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |http|
http.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
line = line.chomp
if line.start_with?("event: ") then event = line[7..]
elsif line.start_with?("data: ") && event == "delta"
raw << (JSON.parse(line[6..])["text"] || "")
# "terms", "ambiguities", "coverage", "context_md" advance the stage.
end
# `tick` events are heartbeats; a browser gets those instead of deltas.
end
end
end
end
result = JSON.parse(raw[raw.index("{")..raw.rindex("}")])
puts "#{result['verdict']} #{result['terms'].length} terms"
<?php
// Server-sent events: the result arrives in chunks, so a UI can show progress.
$raw = "";
$event = null;
$ch = curl_init(BASE . "/run-stream");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($input));
curl_setopt($ch, CURLOPT_HTTPHEADER, [
"Authorization: Bearer " . TOKEN,
"Content-Type: application/json",
"Idempotency-Key: " . $key,
"Accept: text/event-stream",
]);
curl_setopt($ch, CURLOPT_WRITEFUNCTION, function ($ch, $chunk) use (&$raw, &$event) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, "event: ")) {
$event = substr($line, 7); // "delta", "done" - or "tick" in a browser
} elseif (str_starts_with($line, "data: ") && $event === "delta") {
$raw .= json_decode(substr($line, 6), true)["text"] ?? "";
}
}
return strlen($chunk);
});
curl_exec($ch);
curl_close($ch);
$result = json_decode(substr($raw, strpos($raw, "{")), true);
echo $result["verdict"], " ", count($result["terms"]), " terms", PHP_EOL;
// Server-sent events: the result arrives in chunks, so a UI can show progress.
var stream = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run-stream");
stream.Headers.Add("Authorization", "Bearer aut_YOUR_TOKEN");
stream.Headers.Add("Idempotency-Key", key);
stream.Headers.Add("Accept", "text/event-stream");
stream.Content = JsonContent.Create(input);
using var res = await Http.SendAsync(stream, HttpCompletionOption.ResponseHeadersRead);
using var reader = new StreamReader(await res.Content.ReadAsStreamAsync());
var raw = new StringBuilder();
string? evt = null;
while (await reader.ReadLineAsync() is { } line)
{
if (line.StartsWith("event: ")) evt = line[7..]; // delta, done - or tick in a browser
else if (line.StartsWith("data: ") && evt == "delta")
{
var d = JsonSerializer.Deserialize<JsonElement>(line[6..]);
if (d.TryGetProperty("text", out var t)) raw.Append(t.GetString());
// Watch raw for "terms", "ambiguities", "coverage", "context_md".
}
}
Console.WriteLine(raw.ToString());
7. Parse the result
data.output.output is a string holding one JSON object — unwrap
twice. The web app strips an optional code fence, takes everything from the first { to
the last }, parses that, and then normalizes it. Doing the same two things — the slice
and the normalization — is what makes a caller robust against the small variations a model
produces.
Here is an abbreviated glossary reply for the billing PRD above, structurally complete:
{
"lane": "glossary",
"title": "Subscription billing - PRD v0.3",
"summary": "The PRD describes plan changes, trials, cancellation, pausing and per-seat pricing for a three-tier B2B subscription. The one thing a reader must know is that cancellation is stated two incompatible ways, and every other rule about access depends on which one is true.",
"context_name": "Subscription Billing",
"context_summary": "The context that owns plans, subscriptions, invoices and seats for a paying organisation.",
"verdict": "contradictory",
"verdict_reason": "Cancellation is defined twice with opposite consequences for access, so no rule about end-of-access can be implemented from this document.",
"terms": [
{ "id": "T-001", "term": "Subscription", "definition": "One organisation's paid relationship to a tier, billed per period and per seat.",
"avoid": ["package"], "group": "Core", "evidence": "A user picks a package at sign-up and can change tier at any time",
"confidence": "firm" },
{ "id": "T-002", "term": "Seat", "definition": "An invited member who has accepted the invitation; the unit Team and Business are priced by.",
"avoid": [], "group": "Pricing", "evidence": "A seat is an invited member who has accepted; pending invitations do not count",
"confidence": "firm" },
{ "id": "T-003", "term": "Customer", "definition": "The organisation that holds the subscription and receives the invoice.",
"avoid": ["client", "user"], "group": "Core", "evidence": "A client can cancel from the billing page",
"confidence": "tentative" }
],
"ambiguities": [
{ "id": "A-001", "term": "Cancellation", "kind": "contradiction",
"summary": "Cancellation ends access immediately in one section and at the end of the paid period in another.",
"evidence": ["Cancellation takes effect immediately: access ends and no further invoices are raised",
"the customer keeps read-only access until the end of the period they already paid for"],
"proposal": "Name the two cases separately: voluntary cancellation ends access at once; cancellation after failed payment leaves read-only access until the paid period ends.",
"decidable_from_text": false, "owner_hint": "Head of Billing" },
{ "id": "A-002", "term": "Customer", "kind": "synonyms",
"summary": "customer, user and client are used interchangeably for the party that holds the subscription.",
"evidence": ["A user picks a package at sign-up", "A client can cancel from the billing page"],
"proposal": "Use Customer for the paying organisation and Member for a person inside it.",
"decidable_from_text": true, "owner_hint": "" }
],
"excluded": [
{ "term": "billing page", "reason": "A screen in the product, not a concept of the billing domain." }
],
"coverage": [
{ "id": "C-001", "term": "customer", "status": "defined", "ref": "T-003" },
{ "id": "C-002", "term": "seat", "status": "defined", "ref": "T-002" },
{ "id": "C-003", "term": "billing page", "status": "excluded", "ref": "" }
],
"credential_seen": false,
"notes_on_input": "",
"context_md": "# Subscription Billing\n\nThe context that owns plans, subscriptions, invoices and seats...\n\n## Terms\n\n**Subscription**: One organisation's paid relationship to a tier...\n\n**Seat**: An invited member who has accepted...\n\n**Customer**: The organisation that holds the subscription...\n\n## Avoid\n\n- package (use Subscription)\n- client, user (use Customer)\n"
}
And the questionnaire reply for the same document, aimed at Maya:
{
"lane": "questionnaire",
"title": "Questions for the Head of Billing",
"summary": "Two decisions block the plan-change feature: what cancellation does to access, and which word names the paying party.",
"purpose": "Settle the cancellation rule and the naming of the paying party so the PRD can be implemented.",
"from_line": "From: the PM on the plan-change feature",
"to_line": "To: Maya, Head of Billing",
"use_line": "Your answers go straight into PRD v0.4 and into the glossary the team builds from.",
"context_paragraph": "The PRD lets a customer change tier without a support ticket. Two rules in it point in opposite directions, and both are yours to decide.",
"how_to_answer": "Answer under each question. One or two sentences is plenty; where you are unsure, say so and name who is.",
"themes": [
{ "heading": "Cancellation and access",
"questions": [
{ "id": "Q-001", "question": "When a customer cancels voluntarily, does access end that moment or at the end of the period they have paid for?",
"why": "The PRD states both, and every downstream rule about invoices and read-only access depends on the answer.",
"covers": ["N-001"], "priority": 1 },
{ "id": "Q-002", "question": "Should cancellation after three failed payments behave differently from a cancellation the customer chose?",
"why": "The failed-payment section grants read-only access that the voluntary path denies.",
"covers": ["N-001"], "priority": 1 }
] },
{ "heading": "What we call the paying party",
"questions": [
{ "id": "Q-003", "question": "Which single word should invoices and the billing page use for the party that pays: customer, client or account?",
"why": "Three words appear in one document, and the invoice template has to pick one.",
"covers": ["N-002"], "priority": 2 }
] }
],
"closing": "Anything else about billing that you expect to change this quarter and that I should not design around?",
"coverage": [
{ "id": "N-001", "need": "Whether a cancellation ends access immediately or at the end of the paid period.",
"status": "covered", "question_ids": ["Q-001", "Q-002"], "note": "Split into the voluntary and the failed-payment case." },
{ "id": "N-002", "need": "Whether customer, user and client are one thing or several.",
"status": "partial", "question_ids": ["Q-003"],
"note": "The question settles the invoice wording; whether user is a different concept is answerable from the document itself." }
],
"credential_seen": false,
"notes_on_input": "",
"questionnaire_md": "# Questions for the Head of Billing\n\n**To:** Maya, Head of Billing\n\n## Cancellation and access\n\n### When a customer cancels voluntarily, does access end that moment or at the end of the period they have paid for?\n\n>\n\n"
}
What the normalizer does to it
The web app does not trust the reply verbatim, and neither should a caller. These are the behaviours you will actually hit:
- A
lanethat is neitherglossarynorquestionnaireis inferred:questionnaireif the reply hasthemesorquestionnaire_md, otherwiseglossary— and a sentence is appended tonotes_on_inputsaying the reply did not name its lane. Branch on the normalizedlane. - Ids are re-sequenced positionally —
T-001,T-002, … ,A-001, … , andQ-001upward across themes, so question numbering is continuous rather than per-theme. A renderer keyed onT-004finds it even when the model skippedT-003. It also meanscoverage[].refandcovers[]must be validated against the ids you hold after normalization, not the ones the model wrote. - An unrecognised
verdictbecomesdrifting. An unrecognised ambiguitykindbecomesundefined. An unrecognisedconfidencebecomestentative. - An unrecognised glossary
coverage[].statusbecomesexcluded; an unrecognised questionnaire one becomespartial. Aprioritythat is not 1, 2 or 3 becomes2. decidable_from_textis strict: only a realtruecounts, so a missing field reads as "a person has to settle it".- Entries are dropped rather than repaired: a term with neither
termnordefinition, an ambiguity with neithertermnorsummary, anexcludedentry with noterm, acoverageentry with noid, a theme with no questions, a question with noquestiontext. - If a glossary reply has zero terms and zero ambiguities, or a questionnaire reply has no themes and no
questionnaire_md, parsing throws. That is the awkward case:/runsucceeded, you were charged, and the client still rejects the reply. Handle it as a retry with aretry_noteand a bumped attempt suffix, not as a transport error — the web app does exactly one automatic retry and then shows the raw reply. - An empty
titlebecomesUntitled. Strings are coerced, so a number where a string was expected becomes its text rather than an error.
The verdict rule
The glossary verdict is not free-form, and it is worth re-deriving rather than trusting:
contradictory if any ambiguity has kind == "contradiction"
drifting else if any ambiguity has kind "synonyms" or "overloaded"
consistent otherwise
The client computes this from the ambiguity list it received and warns when the returned
verdict disagrees, treating the ambiguity list as authoritative. Do
the same: a gate that reads verdict alone can be talked out of failing by a reply that
lists a contradiction and then calls itself drifting.
Invariants worth asserting in CI
- The verdict matches the rule above, recomputed from
ambiguities[].kind. - Every id in
prescan_facts.candidatesappears exactly once incoverage, and no id you did not send appears there. - Every
coverage[].refnames a realT-orA-id in the same reply. - Every word in a term's
avoidlist actually occurs in the document, and is not itself a canonical term somewhere else in the reply. context_mdcarries each term exactly once as**Term**:, and no bold entry that is not a term.- A document of 200 words or more yields at least three terms.
- Questionnaire: every need id appears exactly once in
coverage; everyquestion_idsentry names a real question; everycoversentry names a need you sent; there are between five and twenty questions; no question contains two question marks or runs past 45 words; andquestionnaire_mdcontains an answer stub line — a line that is just>. credential_seenistrueonly when the paste really did carry a secret — and when it is true, the reply repeats no part of the value. Treat it as a signal to rotate, and keep it out of your logs.
# The result JSON is a string inside the envelope, so unwrap it twice.
RESULT=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])')
printf '%s' "$RESULT" | python3 -c '
import sys, json
r = json.load(sys.stdin)
print(r["lane"], "|", r["title"])
print(r["verdict"], "-", r["verdict_reason"])
for t in r["terms"]:
avoid = ", ".join(t["avoid"]) or "-"
print(" %s %-14s %-9s avoid: %s" % (t["id"], t["term"], t["confidence"], avoid))
for a in r["ambiguities"]:
who = "the text settles it" if a["decidable_from_text"] else "needs " + (a["owner_hint"] or "a person")
print(" %s %-14s %-14s %s" % (a["id"], a["term"], a["kind"], who))
'
# Re-derive the verdict from the ambiguities rather than trusting the field.
printf '%s' "$RESULT" | python3 -c '
import sys, json
r = json.load(sys.stdin)
kinds = [a["kind"] for a in r["ambiguities"]]
want = "contradictory" if "contradiction" in kinds else \
"drifting" if ("synonyms" in kinds or "overloaded" in kinds) else "consistent"
if r["verdict"] != want:
raise SystemExit("verdict says %s but the ambiguities make it %s" % (r["verdict"], want))
print("verdict reconciles:", want)
'
# Every prescan candidate id must come back exactly once in coverage.
printf '%s' "$RESULT" | python3 -c '
import sys, json
seen = [c["id"] for c in json.load(sys.stdin)["coverage"]]
want = ["C-001", "C-002", "C-003"]
bad = [i for i in want if seen.count(i) != 1] + [i for i in seen if i not in want]
if bad:
raise SystemExit("coverage drift: " + ", ".join(bad))
print("coverage reconciles")
'
# The CONTEXT.md is ready to commit as-is.
printf '%s' "$RESULT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["context_md"])' > CONTEXT.md
result = json.loads(job["output"]["output"])
assert result["lane"] == "glossary" # branch on the REPLY's lane, not the task you sent
# 1. Re-derive the verdict; the ambiguity list is authoritative.
kinds = [a["kind"] for a in result["ambiguities"]]
want = ("contradictory" if "contradiction" in kinds
else "drifting" if ("synonyms" in kinds or "overloaded" in kinds)
else "consistent")
if result["verdict"] != want:
raise RuntimeError(f"verdict says {result['verdict']} but the ambiguities make it {want}")
# 2. Every candidate id appears exactly once in coverage, and nothing else does.
sent = [c["id"] for c in INPUT["prescan_facts"]["candidates"]]
seen = [c["id"] for c in result["coverage"]]
missing = [i for i in sent if seen.count(i) != 1]
extra = [i for i in seen if i not in sent]
if missing or extra:
raise RuntimeError(f"coverage drift: missing={missing} extra={extra}")
# Sending empty candidates is legitimate - and then `seen` is empty and this
# check is vacuous. That is the trade: no facts in, no reconciliation out.
# 3. Coverage refs and avoid lists must point at things that exist.
ids = {t["id"] for t in result["terms"]} | {a["id"] for a in result["ambiguities"]}
for c in result["coverage"]:
if c["ref"] and c["ref"] not in ids:
raise RuntimeError(f"coverage {c['id']} points at {c['ref']}, which is not in the reply")
canonical = {t["term"].strip().lower() for t in result["terms"]}
doc = INPUT["document"].lower()
for t in result["terms"]:
for a in t["avoid"]:
if a.strip().lower() in canonical:
raise RuntimeError(f"{a} is both a canonical term and an Avoid word under {t['term']}")
if a.strip().lower().rstrip("s") not in doc:
print("note:", a, "is on an Avoid list but does not occur in the document")
# 4. CONTEXT.md lists each term exactly once.
md = result["context_md"]
for t in result["terms"]:
if md.count(f"**{t['term']}**:") != 1:
raise RuntimeError(f"CONTEXT.md does not list {t['term']} exactly once")
# 5. A truncated reply is a prefix, not a result. Retry, do not repair.
if job.get("truncated"):
INPUT["retry_note"] = ("The previous reply was cut short. Return the same terms but with "
"shorter definitions, and keep context_md complete.")
# ... resubmit with an incremented attempt suffix in the Idempotency-Key.
pathlib.Path("CONTEXT.md").write_text(md, encoding="utf-8")
print(result["verdict"], "-", result["verdict_reason"])
for a in result["ambiguities"]:
print(a["id"], a["kind"], a["term"], "decidable" if a["decidable_from_text"] else a["owner_hint"])
const result = JSON.parse(job.output.output);
if (result.lane !== "glossary") throw new Error(`unexpected lane ${result.lane}`);
// 1. Re-derive the verdict; the ambiguity list is authoritative.
const kinds = result.ambiguities.map((a) => a.kind);
const want = kinds.includes("contradiction")
? "contradictory"
: kinds.includes("synonyms") || kinds.includes("overloaded")
? "drifting"
: "consistent";
if (result.verdict !== want) {
throw new Error(`verdict says ${result.verdict} but the ambiguities make it ${want}`);
}
// 2. Every candidate id appears exactly once in coverage, and nothing else does.
const sent = INPUT.prescan_facts.candidates.map((c) => c.id);
const seen = result.coverage.map((c) => c.id);
const missing = sent.filter((id) => seen.filter((s) => s === id).length !== 1);
const extra = seen.filter((id) => !sent.includes(id));
if (missing.length || extra.length) {
throw new Error(`coverage drift: missing=${missing} extra=${extra}`);
}
// 3. Coverage refs point at ids that exist AFTER normalization.
const ids = new Set([...result.terms.map((t) => t.id), ...result.ambiguities.map((a) => a.id)]);
for (const c of result.coverage) {
if (c.ref && !ids.has(c.ref)) throw new Error(`coverage ${c.id} points at ${c.ref}`);
}
// 4. No word is both canonical and on an Avoid list.
const canonical = new Set(result.terms.map((t) => t.term.trim().toLowerCase()));
for (const t of result.terms) {
for (const a of t.avoid) {
if (canonical.has(a.trim().toLowerCase())) {
throw new Error(`${a} is both a canonical term and an Avoid word under ${t.term}`);
}
}
}
// 5. A truncated reply is a prefix, not a result. Retry, do not repair.
if (job.truncated) {
INPUT.retry_note = "The previous reply was cut short. Shorter definitions, complete context_md.";
}
console.log(result.verdict, "-", result.verdict_reason);
for (const t of result.terms) console.log(t.id, t.term, t.confidence, t.avoid.join("/"));
console.log(result.context_md);
type term struct {
ID string `json:"id"`
Term string `json:"term"`
Definition string `json:"definition"`
Avoid []string `json:"avoid"`
Group string `json:"group"`
Evidence string `json:"evidence"`
Confidence string `json:"confidence"`
}
type ambiguity struct {
ID string `json:"id"`
Term string `json:"term"`
Kind string `json:"kind"`
Summary string `json:"summary"`
Evidence []string `json:"evidence"`
Proposal string `json:"proposal"`
DecidableFromText bool `json:"decidable_from_text"`
OwnerHint string `json:"owner_hint"`
}
type glossary struct {
Lane string `json:"lane"`
Title string `json:"title"`
Summary string `json:"summary"`
ContextName string `json:"context_name"`
Verdict string `json:"verdict"`
VerdictReason string `json:"verdict_reason"`
Terms []term `json:"terms"`
Ambiguities []ambiguity `json:"ambiguities"`
Excluded []struct {
Term string `json:"term"`
Reason string `json:"reason"`
} `json:"excluded"`
Coverage []struct {
ID string `json:"id"`
Term string `json:"term"`
Status string `json:"status"`
Ref string `json:"ref"`
} `json:"coverage"`
CredentialSeen bool `json:"credential_seen"`
NotesOnInput string `json:"notes_on_input"`
ContextMD string `json:"context_md"`
}
var g glossary
if err := json.Unmarshal([]byte(job.Output.Output), &g); err != nil {
panic(err)
}
// Re-derive the verdict from the ambiguities rather than trusting the field.
want := "consistent"
for _, a := range g.Ambiguities {
if a.Kind == "synonyms" || a.Kind == "overloaded" {
want = "drifting"
}
}
for _, a := range g.Ambiguities {
if a.Kind == "contradiction" {
want = "contradictory"
}
}
if g.Verdict != want {
panic("verdict says " + g.Verdict + " but the ambiguities make it " + want)
}
// Every prescan candidate id must come back exactly once in coverage.
count := map[string]int{}
for _, c := range g.Coverage {
count[c.ID]++
}
for _, id := range []string{"C-001", "C-002"} {
if count[id] != 1 {
panic("unreconciled prescan candidate: " + id)
}
}
_ = os.WriteFile("CONTEXT.md", []byte(g.ContextMD), 0o644)
fmt.Println(g.Verdict, len(g.Terms), "terms,", len(g.Ambiguities), "ambiguities")
// The result JSON is a string inside data.output.output - parse it, then check
// the invariants before you trust it:
//
// 1. re-derive the verdict: "contradictory" if any ambiguity kind is
// "contradiction", else "drifting" if any is "synonyms" or "overloaded",
// else "consistent" - and compare it with the verdict field;
// 2. every prescan_facts.candidates id appears exactly once in coverage, and
// no id you did not send appears there;
// 3. every coverage[].ref names a T- or A- id present after normalization,
// which re-sequences ids positionally;
// 4. no word is both a canonical term and on another term's avoid list;
// 5. context_md lists each term exactly once as **Term**:.
//
// A `truncated` job is a prefix, not a result: resubmit with a retry_note such
// as "The previous reply was cut short. Shorter definitions, complete
// context_md." and an incremented attempt suffix on the Idempotency-Key.
String resultJson = /* data.output.output */ call("jobs/" + jobId, null);
System.out.println(resultJson);
// Branch on the reply's own "lane" field: an unrecognised task degrades to the
// closest lane and says so in notes_on_input rather than failing.
result = JSON.parse(job["output"]["output"])
abort "unexpected lane #{result['lane']}" unless result["lane"] == "glossary"
# 1. Re-derive the verdict; the ambiguity list is authoritative.
kinds = result["ambiguities"].map { |a| a["kind"] }
want = if kinds.include?("contradiction") then "contradictory"
elsif kinds.include?("synonyms") || kinds.include?("overloaded") then "drifting"
else "consistent"
end
raise "verdict says #{result['verdict']} but the ambiguities make it #{want}" if result["verdict"] != want
# 2. Every candidate id appears exactly once in coverage.
sent = input["prescan_facts"]["candidates"].map { |c| c["id"] }
seen = result["coverage"].map { |c| c["id"] }
missing = sent.reject { |id| seen.count(id) == 1 }
extra = seen - sent
raise "coverage drift: #{missing} / #{extra}" unless missing.empty? && extra.empty?
# 3. Coverage refs and avoid lists must point at things that exist.
ids = (result["terms"] + result["ambiguities"]).map { |x| x["id"] }
result["coverage"].each do |c|
raise "coverage #{c['id']} points at #{c['ref']}" if !c["ref"].empty? && !ids.include?(c["ref"])
end
canonical = result["terms"].map { |t| t["term"].strip.downcase }
result["terms"].each do |t|
t["avoid"].each do |a|
raise "#{a} is both canonical and an Avoid word" if canonical.include?(a.strip.downcase)
end
end
File.write("CONTEXT.md", result["context_md"])
puts "#{result['verdict']} - #{result['verdict_reason']}"
result["ambiguities"].each { |a| puts "#{a['id']} #{a['kind']} #{a['term']}" }
<?php
$result = json_decode($job["output"]["output"], true);
if ($result["lane"] !== "glossary") { throw new RuntimeException("unexpected lane"); }
// 1. Re-derive the verdict; the ambiguity list is authoritative.
$kinds = array_column($result["ambiguities"], "kind");
$want = in_array("contradiction", $kinds, true) ? "contradictory"
: (array_intersect(["synonyms", "overloaded"], $kinds) ? "drifting" : "consistent");
if ($result["verdict"] !== $want) {
throw new RuntimeException("verdict says {$result['verdict']} but the ambiguities make it {$want}");
}
// 2. Every candidate id appears exactly once in coverage.
$sent = array_column($input["prescan_facts"]["candidates"], "id");
$seen = array_column($result["coverage"], "id");
$counts = array_count_values($seen);
foreach ($sent as $id) {
if (($counts[$id] ?? 0) !== 1) {
throw new RuntimeException("unreconciled prescan candidate: " . $id);
}
}
foreach (array_diff($seen, $sent) as $id) {
throw new RuntimeException("coverage names " . $id . ", which was not sent");
}
// 3. No word is both canonical and on an Avoid list.
$canonical = array_map(fn($t) => strtolower(trim($t["term"])), $result["terms"]);
foreach ($result["terms"] as $t) {
foreach ($t["avoid"] as $a) {
if (in_array(strtolower(trim($a)), $canonical, true)) {
throw new RuntimeException("{$a} is both a canonical term and an Avoid word");
}
}
}
file_put_contents("CONTEXT.md", $result["context_md"]);
echo $result["verdict"], " - ", $result["verdict_reason"], PHP_EOL;
var result = JsonSerializer.Deserialize<JsonElement>(resultJson);
// 1. Re-derive the verdict; the ambiguity list is authoritative.
var kinds = result.GetProperty("ambiguities").EnumerateArray()
.Select(a => a.GetProperty("kind").GetString())
.ToList();
var want = kinds.Contains("contradiction") ? "contradictory"
: kinds.Contains("synonyms") || kinds.Contains("overloaded") ? "drifting"
: "consistent";
if (result.GetProperty("verdict").GetString() != want)
throw new Exception($"verdict says {result.GetProperty("verdict")} but the ambiguities make it {want}");
// 2. Every prescan candidate id appears exactly once in coverage.
var seen = result.GetProperty("coverage").EnumerateArray()
.Select(c => c.GetProperty("id").GetString())
.ToList();
foreach (var id in new[] { "C-001", "C-002" })
{
if (seen.Count(s => s == id) != 1)
throw new Exception($"unreconciled prescan candidate: {id}");
}
// 3. No word is both canonical and on an Avoid list.
var terms = result.GetProperty("terms").EnumerateArray().ToList();
var canonical = terms.Select(t => t.GetProperty("term").GetString()!.Trim().ToLowerInvariant()).ToHashSet();
foreach (var t in terms)
foreach (var a in t.GetProperty("avoid").EnumerateArray())
if (canonical.Contains(a.GetString()!.Trim().ToLowerInvariant()))
throw new Exception($"{a} is both a canonical term and an Avoid word");
File.WriteAllText("CONTEXT.md", result.GetProperty("context_md").GetString());
foreach (var t in terms)
Console.WriteLine($"{t.GetProperty("id")} {t.GetProperty("term")} {t.GetProperty("confidence")}");
The output contract
Every key in the object, as the web app reads it. First the envelope both lanes share:
| key | type | meaning |
|---|---|---|
lane | enum | glossary or questionnaire — which contract the rest of the object follows. Inferred from the fields present when the model omits it or invents a value, and the inference is recorded in notes_on_input. Branch on this, not on the task you sent. |
title | string | A short name for this run, taken from the document's own subject. Empty becomes Untitled. |
summary | string | Two to four sentences: what the document is about and the one thing the reader must know. |
coverage | object[] | The reconciliation table. Shape differs per lane — see below. One entry per prescan_facts.candidates id in the glossary lane, one per needs id in the questionnaire lane. |
credential_seen | boolean | true when the paste looked like it carried a password, API key, token or private key. The reply then repeats no part of the value anywhere and says in notes_on_input that it should be rotated. |
notes_on_input | string | "" when there is nothing to say. Carries: that the middle of the document was clipped, that a decisions entry contradicts the document, that the task was missing or unrecognised and which lane was chosen instead. |
Glossary body
| key | type | meaning |
|---|---|---|
context_name | string | The bounded context the document is about, in two or three words — "Subscription Billing". Empty falls back to title. |
context_summary | string | One or two sentences saying what this context owns and where its edge is. |
verdict | enum | consistent, drifting or contradictory. Anything else normalizes to drifting. The single value a CI gate should branch on — after re-deriving it from the ambiguity kinds. |
verdict_reason | string | One sentence naming the ambiguity that decided the verdict. |
terms | object[] | {id, term, definition, avoid, group, evidence, confidence}. Ids are sequential T-001, T-002, … definition is one or two sentences; avoid is the words the document uses for the same thing that the team should stop using; group is a heading such as Core or Pricing, defaulting to General; evidence is a verbatim quote from the document, trimmed to about 200 characters. |
ambiguities | object[] | {id, term, kind, summary, evidence, proposal, decidable_from_text, owner_hint}. Ids are sequential A-001, … evidence is an array of verbatim quotes; proposal is the resolution the model would pick; decidable_from_text says whether the document itself settles it, and when it is false, owner_hint names the kind of person who can. Those two fields are exactly what the questionnaire lane consumes. |
excluded | object[] | {term, reason} — words considered and deliberately left out, because they are general programming, business or project-management vocabulary the document merely uses rather than concepts of this context. Shipping the rejects is what makes the term list auditable. |
coverage | object[] | {id, term, status, ref}. One entry per prescan candidate id, exactly once, with status in defined, avoid, excluded, ambiguity, not_a_term, and ref pointing at the T- or A- id that handled it ("" otherwise). |
context_md | string | A whole CONTEXT.md in one JSON string, ready to commit at the repository root: the context name, the summary, each term as **Term**: definition exactly once, and an Avoid section. "" when the document gives nothing to base it on. |
Questionnaire body
| key | type | meaning |
|---|---|---|
purpose | string | One sentence saying what the answers will be used for. |
from_line / to_line | string | Who is asking and who is being asked, drawn from recipient. |
use_line | string | What happens to the answers — the sentence that makes the recipient's time feel spent rather than taken. |
context_paragraph | string | Enough of the document for the recipient to answer without reading it. |
how_to_answer | string | The instruction line: answer under each question, one or two sentences, say when you are unsure. |
themes | object[] | {heading, questions}, where each question is {id, question, why, covers, priority}. Question ids run Q-001 upward across themes, not per theme. why says what changes depending on the answer; covers lists the N- need ids the question serves; priority is 1, 2 or 3, most important first. |
closing | string | The catch-all question at the end. |
coverage | object[] | {id, need, status, question_ids, note}. One entry per need you sent, exactly once. status is covered, partial, answered_in_document — the document already settles it, so no question was spent on it — or not_covered. question_ids names the questions that serve it. |
questionnaire_md | string | The whole questionnaire as Markdown in one JSON string, with each question as an ### heading and a blockquote answer stub — a line that is just > — under it, ready to paste into an email or a doc. |
The enums
| field | values | notes |
|---|---|---|
verdict | consistent, drifting, contradictory | consistent: the document names one thing one way. drifting: it has competing names or an overloaded word, but nothing that makes two rules incompatible. contradictory: at least one place where following the document two ways gives two different systems. Derived, not chosen — see the verdict rule. Unrecognised values normalize to drifting. |
ambiguities[].kind | overloaded, synonyms, undefined, contradiction, boundary | overloaded: one word carries two meanings. synonyms: several words carry one meaning. undefined: a word the document leans on but never defines. contradiction: two statements that cannot both hold. boundary: it is unclear whether a concept belongs to this context at all. Unrecognised values normalize to undefined. |
terms[].confidence | firm, tentative | firm means the document (or an authoritative decisions entry) states the definition; tentative means it was inferred from usage and should be confirmed. Unrecognised values normalize to tentative, which is the safe direction. |
coverage[].status (glossary) | defined, avoid, excluded, ambiguity, not_a_term | What became of a scanned candidate. not_a_term is a legitimate answer — the scanner is mechanical and is allowed to be wrong. Unrecognised values normalize to excluded. |
coverage[].status (questionnaire) | covered, partial, answered_in_document, not_covered | answered_in_document means no question was spent on it because the text already settles it. Unrecognised values normalize to partial. |
questions[].priority | 1, 2, 3 | A number, not a string. 1 is "the project stops without this". Anything else normalizes to 2. |
redefinitions[].kind (input) | defined_twice, conflicting_rules | Sent by you inside prescan_facts, not returned. |
The reply never echoes a secret value. If the paste contains an API key, a password in a
connection string or a token in a log line, credential_seen comes back
true and the note says to rotate it — the value itself appears in no term, no quote
and no context_md.
8. Use it in CI
The worked example: a job reads the spec out of the repository, extracts its glossary, writes
CONTEXT.md, and exits non-zero when the verdict is contradictory — the
document says two incompatible things and nobody can implement it as written. Fail on
contradictory; report on drifting, which is normal in a living
document and would otherwise make the gate noise people learn to ignore. Derive the
Idempotency-Key from the spec's contents so a re-run of the same commit replays the same job
instead of re-billing, and only bump the attempt suffix when the text actually changed.
The script sends prescan_facts with empty arrays, which is the honest thing for a
caller with no local scanner — and it means coverage comes back empty, so the gate
leans on the verdict and the ambiguity list instead of on reconciliation. If you do have ids to
send, send them: the reconciliation check is the strongest signal in the reply.
#!/bin/sh
# glossary-gate.sh - fail the build when the spec contradicts itself.
set -eu
BASE="https://api.skillsafe.ai/v1/app-api"
TOKEN="$SKILLSAFE_TOKEN" # from https://glossary-desk.skillsafe.ai/tokens.html
SPEC="${1:-docs/spec.md}"
[ -f "$SPEC" ] || { echo "glossary-desk: no $SPEC in this repository"; exit 0; }
# 1. Build the input. An API caller may send empty prescan facts - and then the
# coverage list comes back empty, so the gate reads the verdict instead.
INPUT=$(SPEC="$SPEC" python3 -c '
import json, os, pathlib
text = pathlib.Path(os.environ["SPEC"]).read_text(encoding="utf-8")
if len(text) > 60000: # clip the MIDDLE, keep both ends
head, tail = text[:37000], text[-20000:]
cut = len(text) - len(head) - len(tail)
text = head + "\n\n[... %d characters cut from the middle of the document - the beginning and the end are kept ...]\n\n" % cut + tail
print(json.dumps({
"task": "glossary",
"document": text,
"context_hint": "CI gate on every pull request that touches the spec.",
"decisions": "",
"prescan_facts": {"stats": {}, "candidates": [], "clusters": [],
"redefinitions": [], "clipped": {"cut": 0}},
}))')
KEY="glossary-desk:glossary:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" -H "Idempotency-Key: $KEY" \
-d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
OUT=$(curl -sS "$BASE/jobs/$JOB" -H "Authorization: Bearer $TOKEN")
STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
sleep 2
done
# 2. Gate on the verdict, re-derived from the ambiguities. Fail on
# contradictory; report drifting without failing.
printf '%s' "$OUT" | python3 -c '
import sys, json, pathlib
job = json.load(sys.stdin)["data"]
if job.get("truncated"):
raise SystemExit("::error::glossary-desk: reply was cut short, the glossary is incomplete")
r = json.loads(job["output"]["output"])
kinds = [a["kind"] for a in r["ambiguities"]]
verdict = "contradictory" if "contradiction" in kinds else \
"drifting" if ("synonyms" in kinds or "overloaded" in kinds) else "consistent"
pathlib.Path("CONTEXT.md").write_text(r["context_md"], encoding="utf-8")
for a in r["ambiguities"]:
who = "the text settles it" if a["decidable_from_text"] else (a["owner_hint"] or "a person")
print("%s %-14s %-14s %s [%s]" % (a["id"], a["kind"], a["term"], a["summary"], who))
print(verdict, "-", r["verdict_reason"])
if verdict == "contradictory":
blocking = [a["id"] for a in r["ambiguities"] if a["kind"] == "contradiction"]
raise SystemExit("::error::glossary-desk: the spec contradicts itself (%s)" % ",".join(blocking))
if verdict == "drifting":
print("::warning::glossary-desk: the spec is drifting - competing names, but nothing blocking")
'
#!/usr/bin/env python3
"""glossary_gate.py - fail CI when the spec contradicts itself.
Reuses the `call` helper from section 2. Exits 1 on a contradictory verdict and
on a truncated reply, which is a prefix and not a glossary. A drifting verdict
is reported and does not fail: it is the normal state of a living document.
"""
import hashlib, json, pathlib, sys, time, urllib.request
SPEC = pathlib.Path(sys.argv[1] if len(sys.argv) > 1 else "docs/spec.md")
if not SPEC.is_file():
sys.exit(f"glossary-desk: no {SPEC} in this repository")
text = SPEC.read_text(encoding="utf-8")
if len(text) > 60000: # clip the MIDDLE, keep both ends
head, tail = text[:37000], text[-20000:]
cut = len(text) - len(head) - len(tail)
text = (head + f"\n\n[... {cut} characters cut from the middle of the document - "
"the beginning and the end are kept ...]\n\n" + tail)
INPUT = {
"task": "glossary",
"document": text,
"context_hint": "CI gate on every pull request that touches the spec.",
"decisions": "",
# Empty prescan facts are legitimate. The cost is honest: `coverage` comes
# back empty and there is nothing to reconcile, so this gate reads the
# verdict and the ambiguity list instead.
"prescan_facts": {"stats": {}, "candidates": [], "clusters": [],
"redefinitions": [], "clipped": {"cut": 0}},
}
digest = hashlib.sha256(json.dumps(INPUT, sort_keys=True).encode()).hexdigest()[:16]
key = f"glossary-desk:glossary:{digest}:a1"
req = urllib.request.Request(f"{BASE}/run", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
with urllib.request.urlopen(req) as r:
job_id = json.load(r)["data"]["job_id"]
while True:
job = call(f"jobs/{job_id}")
if job["status"] in ("succeeded", "failed"):
break
time.sleep(2)
if job["status"] == "failed":
sys.exit(f"glossary-desk: run failed: {job.get('error')}")
if job.get("truncated"):
sys.exit("glossary-desk: reply was cut short - what came back is a prefix, not a glossary")
r = json.loads(job["output"]["output"])
kinds = [a["kind"] for a in r["ambiguities"]]
verdict = ("contradictory" if "contradiction" in kinds
else "drifting" if ("synonyms" in kinds or "overloaded" in kinds)
else "consistent")
pathlib.Path("CONTEXT.md").write_text(r["context_md"], encoding="utf-8")
for t in r["terms"]:
print(f" {t['id']} {t['term']:20} {t['confidence']:9} avoid: {', '.join(t['avoid']) or '-'}")
for a in r["ambiguities"]:
who = "the text settles it" if a["decidable_from_text"] else (a["owner_hint"] or "a person")
print(f" {a['id']} {a['kind']:14} {a['term']:20} [{who}] {a['summary']}")
print(verdict, "-", r["verdict_reason"], "| charged", job.get("charged_credits"))
if verdict == "contradictory":
blocking = [a["id"] for a in r["ambiguities"] if a["kind"] == "contradiction"]
sys.exit(f"glossary-desk: the spec contradicts itself ({','.join(blocking)})")
if verdict == "drifting":
print("glossary-desk: drifting - competing names, but nothing that blocks implementation")
// glossary-gate.mjs - fail CI when the spec contradicts itself.
// Reuses the `call` helper from section 2.
import { createHash } from "node:crypto";
import { existsSync, readFileSync, writeFileSync } from "node:fs";
const SPEC = process.argv[2] ?? "docs/spec.md";
if (!existsSync(SPEC)) {
console.log(`glossary-desk: no ${SPEC} in this repository`);
process.exit(0);
}
let text = readFileSync(SPEC, "utf8");
if (text.length > 60000) { // clip the MIDDLE, keep both ends
const head = text.slice(0, 37000);
const tail = text.slice(-20000);
const cut = text.length - head.length - tail.length;
text = `${head}\n\n[... ${cut} characters cut from the middle of the document - the beginning and the end are kept ...]\n\n${tail}`;
}
const INPUT = {
task: "glossary",
document: text,
context_hint: "CI gate on every pull request that touches the spec.",
decisions: "",
// Empty prescan facts are legitimate; coverage then comes back empty.
prescan_facts: { stats: {}, candidates: [], clusters: [], redefinitions: [], clipped: { cut: 0 } },
};
const digest = createHash("sha256").update(JSON.stringify(INPUT)).digest("hex").slice(0, 16);
const started = await fetch(`${BASE}/run`, {
method: "POST",
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": `glossary-desk:glossary:${digest}:a1`,
},
body: JSON.stringify(INPUT),
}).then((r) => r.json());
let job = started.data;
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`jobs/${job.job_id}`);
}
if (job.status === "failed") throw new Error("glossary-desk: run failed");
if (job.truncated) throw new Error("glossary-desk: reply was cut short - that is a prefix, not a glossary");
const r = JSON.parse(job.output.output);
const kinds = r.ambiguities.map((a) => a.kind);
const verdict = kinds.includes("contradiction")
? "contradictory"
: kinds.includes("synonyms") || kinds.includes("overloaded")
? "drifting"
: "consistent";
writeFileSync("CONTEXT.md", r.context_md);
for (const a of r.ambiguities) {
const who = a.decidable_from_text ? "the text settles it" : a.owner_hint || "a person";
console.log(` ${a.id} ${a.kind} ${a.term} [${who}] ${a.summary}`);
}
console.log(verdict, "-", r.verdict_reason, "| charged", job.charged_credits);
if (verdict === "contradictory") {
const blocking = r.ambiguities.filter((a) => a.kind === "contradiction").map((a) => a.id);
console.error(`glossary-desk: the spec contradicts itself (${blocking.join(",")})`);
process.exitCode = 1;
} else if (verdict === "drifting") {
console.warn("glossary-desk: drifting - competing names, but nothing blocking");
}
// The gate, on top of the client and the glossary struct from the earlier
// sections: read the spec, extract its glossary, write CONTEXT.md, and exit
// non-zero when the document contradicts itself.
specPath := "docs/spec.md"
if len(os.Args) > 1 {
specPath = os.Args[1]
}
b, err := os.ReadFile(specPath)
if err != nil {
fmt.Println("glossary-desk: no", specPath, "in this repository")
os.Exit(0)
}
text := string(b)
if len(text) > 60000 { // clip the MIDDLE, keep both ends
head, tail := text[:37000], text[len(text)-20000:]
text = fmt.Sprintf("%s\n\n[... %d characters cut from the middle of the document - "+
"the beginning and the end are kept ...]\n\n%s", head, len(text)-57000, tail)
}
input := map[string]any{
"task": "glossary",
"document": text,
"context_hint": "CI gate on every pull request that touches the spec.",
"decisions": "",
// Empty prescan facts are legitimate; coverage then comes back empty.
"prescan_facts": map[string]any{"stats": map[string]any{}, "candidates": []any{},
"clusters": []any{}, "redefinitions": []any{}, "clipped": map[string]any{"cut": 0}},
}
// ... POST /run with the Idempotency-Key, poll jobs/{job_id}, unmarshal into `g`.
verdict := "consistent"
var blocking []string
for _, a := range g.Ambiguities {
if a.Kind == "synonyms" || a.Kind == "overloaded" {
verdict = "drifting"
}
}
for _, a := range g.Ambiguities {
if a.Kind == "contradiction" {
verdict = "contradictory"
blocking = append(blocking, a.ID)
}
}
_ = os.WriteFile("CONTEXT.md", []byte(g.ContextMD), 0o644)
fmt.Println(verdict, "-", g.VerdictReason)
if verdict == "contradictory" {
fmt.Fprintf(os.Stderr, "glossary-desk: the spec contradicts itself (%s)\n", strings.Join(blocking, ","))
os.Exit(1)
}
if verdict == "drifting" {
fmt.Println("glossary-desk: drifting - competing names, but nothing blocking")
}
// The gate, on top of the GlossaryDesk client from section 2. Read the spec into
// `document`, POST /run with the Idempotency-Key, poll jobs/{job_id}, then:
//
// var kinds = /* ambiguities[].kind */;
// String verdict = kinds.contains("contradiction") ? "contradictory"
// : kinds.contains("synonyms") || kinds.contains("overloaded")
// ? "drifting" : "consistent";
// if ("contradictory".equals(verdict)) System.exit(1); // fail the build
// if ("drifting".equals(verdict)) System.out.println("warning: drifting");
//
// Send "prescan_facts": {"stats": {}, "candidates": [], "clusters": [],
// "redefinitions": [], "clipped": {"cut": 0}} when you have no local scanner -
// it is legitimate, and the cost is that `coverage` comes back empty, so the
// gate leans on the verdict and the ambiguity list. A truncated job is also a
// failure: what you hold is a prefix.
var spec = java.nio.file.Path.of(args.length > 0 ? args[0] : "docs/spec.md");
if (!java.nio.file.Files.isRegularFile(spec)) {
System.out.println("glossary-desk: no " + spec + " in this repository");
return;
}
String text = java.nio.file.Files.readString(spec);
if (text.length() > 60000) { // clip the MIDDLE, keep both ends, keep the marker
int cut = text.length() - 57000;
text = text.substring(0, 37000)
+ "\n\n[... " + cut + " characters cut from the middle of the document - "
+ "the beginning and the end are kept ...]\n\n"
+ text.substring(text.length() - 20000);
}
System.out.println(text.length() + " characters of spec to work on");
// Write the reply's context_md to CONTEXT.md once the job succeeds.
# glossary_gate.rb - fail CI when the spec contradicts itself.
# Reuses the `call` helper from section 2.
require "digest"
spec = ARGV[0] || "docs/spec.md"
unless File.file?(spec)
puts "glossary-desk: no #{spec} in this repository"
exit 0
end
text = File.read(spec)
if text.length > 60_000 # clip the MIDDLE, keep both ends
cut = text.length - 57_000
text = text[0, 37_000] +
"\n\n[... #{cut} characters cut from the middle of the document - " \
"the beginning and the end are kept ...]\n\n" + text[-20_000..]
end
input = {
"task" => "glossary",
"document" => text,
"context_hint" => "CI gate on every pull request that touches the spec.",
"decisions" => "",
# Empty prescan facts are legitimate; coverage then comes back empty.
"prescan_facts" => { "stats" => {}, "candidates" => [], "clusters" => [],
"redefinitions" => [], "clipped" => { "cut" => 0 } }
}
key = "glossary-desk:glossary:#{Digest::SHA256.hexdigest(JSON.generate(input))[0, 16]}:a1"
# ... POST /run with that Idempotency-Key, then poll jobs/{job_id} as in section 5.
abort "glossary-desk: reply was cut short - that is a prefix, not a glossary" if job["truncated"]
r = JSON.parse(job["output"]["output"])
kinds = r["ambiguities"].map { |a| a["kind"] }
verdict = if kinds.include?("contradiction") then "contradictory"
elsif kinds.include?("synonyms") || kinds.include?("overloaded") then "drifting"
else "consistent"
end
File.write("CONTEXT.md", r["context_md"])
r["ambiguities"].each { |a| puts " #{a['id']} #{a['kind']} #{a['term']} - #{a['summary']}" }
puts "#{verdict} - #{r['verdict_reason']}"
if verdict == "contradictory"
blocking = r["ambiguities"].select { |a| a["kind"] == "contradiction" }.map { |a| a["id"] }
abort "glossary-desk: the spec contradicts itself (#{blocking.join(',')})"
end
puts "glossary-desk: drifting - competing names, but nothing blocking" if verdict == "drifting"
<?php
// glossary-gate.php - fail CI when the spec contradicts itself.
// Reuses the `call` helper from section 2.
$spec = $argv[1] ?? "docs/spec.md";
if (!is_file($spec)) {
echo "glossary-desk: no {$spec} in this repository", PHP_EOL;
exit(0);
}
$text = file_get_contents($spec);
if (strlen($text) > 60000) { // clip the MIDDLE, keep both ends
$cut = strlen($text) - 57000;
$text = substr($text, 0, 37000)
. "\n\n[... {$cut} characters cut from the middle of the document - "
. "the beginning and the end are kept ...]\n\n"
. substr($text, -20000);
}
$input = [
"task" => "glossary",
"document" => $text,
"context_hint" => "CI gate on every pull request that touches the spec.",
"decisions" => "",
// Empty prescan facts are legitimate; coverage then comes back empty.
"prescan_facts" => ["stats" => (object) [], "candidates" => [], "clusters" => [],
"redefinitions" => [], "clipped" => ["cut" => 0]],
];
$key = "glossary-desk:glossary:" . substr(hash("sha256", json_encode($input)), 0, 16) . ":a1";
// ... POST /run with that Idempotency-Key, then poll jobs/{job_id} as in section 5.
if (!empty($job["truncated"])) {
fwrite(STDERR, "glossary-desk: reply was cut short - that is a prefix\n");
exit(1);
}
$r = json_decode($job["output"]["output"], true);
$kinds = array_column($r["ambiguities"], "kind");
$verdict = in_array("contradiction", $kinds, true) ? "contradictory"
: (array_intersect(["synonyms", "overloaded"], $kinds) ? "drifting" : "consistent");
file_put_contents("CONTEXT.md", $r["context_md"]);
echo $verdict, " - ", $r["verdict_reason"], PHP_EOL;
if ($verdict === "contradictory") {
$blocking = array_column(array_filter($r["ambiguities"], fn($a) => $a["kind"] === "contradiction"), "id");
fwrite(STDERR, "glossary-desk: the spec contradicts itself (" . implode(",", $blocking) . ")\n");
exit(1);
}
if ($verdict === "drifting") {
echo "glossary-desk: drifting - competing names, but nothing blocking", PHP_EOL;
}
// The gate, on top of the GlossaryDesk client from section 2.
var spec = args.Length > 0 ? args[0] : "docs/spec.md";
if (!File.Exists(spec))
{
Console.WriteLine($"glossary-desk: no {spec} in this repository");
return;
}
var text = File.ReadAllText(spec);
if (text.Length > 60000) // clip the MIDDLE, keep both ends, keep the marker
{
var cut = text.Length - 57000;
text = text[..37000]
+ $"\n\n[... {cut} characters cut from the middle of the document - "
+ "the beginning and the end are kept ...]\n\n"
+ text[^20000..];
}
var input = new
{
task = "glossary",
document = text,
context_hint = "CI gate on every pull request that touches the spec.",
decisions = "",
// Empty prescan facts are legitimate; coverage then comes back empty.
prescan_facts = new { stats = new { }, candidates = Array.Empty<object>(),
clusters = Array.Empty<object>(), redefinitions = Array.Empty<object>(),
clipped = new { cut = 0 } },
};
// ... POST /run with the Idempotency-Key, poll jobs/{job_id}, parse data.output.output.
var kinds = result.GetProperty("ambiguities").EnumerateArray()
.Select(a => a.GetProperty("kind").GetString())
.ToList();
var verdict = kinds.Contains("contradiction") ? "contradictory"
: kinds.Contains("synonyms") || kinds.Contains("overloaded") ? "drifting"
: "consistent";
File.WriteAllText("CONTEXT.md", result.GetProperty("context_md").GetString());
Console.WriteLine($"{verdict} - {result.GetProperty("verdict_reason")}");
if (verdict == "contradictory")
{
Console.Error.WriteLine("glossary-desk: the spec contradicts itself");
Environment.Exit(1);
}
if (verdict == "drifting")
Console.WriteLine("glossary-desk: drifting - competing names, but nothing blocking");
Truncation and partial results
When the balance sits between min_credits and hold_credits, the run is
not refused: it executes with a reduced output cap and comes back with truncated: true
on the finished job and on the streaming done event. What you hold then is a prefix of
the reply, not the reply — in the glossary lane the terms may be complete while
coverage, excluded and context_md are missing or cut
mid-string; in the questionnaire lane the themes may be there while the coverage table and
questionnaire_md are not.
Check the flag before you treat a reply as complete, and remember that the client's parser is strict in one direction only: a prefix that still contains at least one term or one ambiguity parses, so a truncated glossary can look like a small glossary. That is why the flag, and not the shape of the object, is the test.
The right response is a retry, not a repair: resubmit with a retry_note asking for
fewer, denser terms and a shorter context_md, and with the attempt suffix on the
Idempotency-Key incremented so the new body is not a replay of the old key. Repairing
truncated JSON by appending closing braces produces something that parses and is not what the model
meant — and in this app it produces a glossary whose coverage silently disagrees with
the terms above it.
One more honest limit: the document is clipped from the middle at 60,000 characters, and
the reply will say so in notes_on_input. A glossary built from a clipped document is a
glossary of the beginning and the end. For anything longer, split the document along its own
section boundaries, run each part, and reconcile the term lists yourself — the terms carry
evidence quotes precisely so that merging two runs is a matter of comparing quotes
rather than trusting two definitions.