← Brainstorm Desk / API
Tokens

Drive Brainstorm Desk from your own code

A review of a shortlist, not a decision. The model never picks the winner: decision is always null, and the decision stays with the decision owner. It never adds an idea that is not in your notes and never invents a rating. A valid register and a ranked matrix establish neither that an idea works nor that it is safe to pursue.

Brainstorm Desk turns a scientific brainstorming session into a structured record and then challenges its shortlist. The two tools of the agent skill @k-dense-ai/scientific-brainstorming (k-dense-ai/scientific-agent-skills) run free in your browser: validate_register.py (the idea register against schema 1.1: errors, warnings and statistics) and evaluate_matrix.py (the scoring sheet against the criteria file: normalised weights, scores, rating-interval and weight-sensitivity ranges, and the shortlist).

Those free tools are not on the API - the page runs them locally. Outside the page, run the skill's own scripts to compute the same facts: python3 scripts/validate_register.py register.json and python3 scripts/evaluate_matrix.py scores.csv --config criteria.json. The API is the two metered lanes, on model gpt-terra: draft turns session notes into a register, a criteria file and a scoring sheet, and review challenges the shortlisted ideas using the facts those tools computed. The natural loop: draft, run the tools, review, revise, re-check, then hand the record to the decision owner.

Two lanes: the task field

task is the one required field and picks the lane. /estimate does not validate the body, so always send a JSON object with task set to one of the two lanes.

taskneedsreturns
draftnotes (the session notes); to revise, also register and criteria (and facts, if you have them)A register (schema 1.1), a criteria file (schema 1.0) and a scoring sheet as CSV, with ratings_source notes (the ratings were in the notes) or blank (an empty sheet to fill), every value the notes did not state in assumptions_made, what you must decide in open_questions and, for a revision, each change in changes.
reviewcriteria + scores + facts; register when you have one; notes as optional contextA status (ready_for_decision_owner, revise_first, not_reviewable), a response to every browser flag, a challenge of every shortlisted idea, fragility notes, safety and feasibility gates, criteria issues, missing perspectives, questions, decision: null, and what nothing here establishes.

Worked example: review

Two shortlisted ideas scored on two criteria: information_gain (higher is better, 1-5, weight 3) and resource_burden (lower is better, 1-5, weight 2). No register is sent, the matrix ranks both, and the page raises one flag: the ideas' rating intervals overlap. Its status is ready_for_decision_owner. Every value is a string: criteria and facts are JSON texts, scores is CSV text.

{
 "task": "review",
 "criteria": "{\"schema_version\":\"1.0\",\"criteria\":[{\"name\":\"information_gain\",\"description\":\"How much the idea would teach us\",\"weight\":3,\"direction\":\"higher\",\"minimum\":1,\"maximum\":5},{\"name\":\"resource_burden\",\"description\":\"Staff time and cost\",\"weight\":2,\"direction\":\"lower\",\"minimum\":1,\"maximum\":5}]}",
 "scores": "idea_id,information_gain,resource_burden\nI1,4,2\nI2,4,3\n",
 "facts": "{\"page\":\"brainstorm-desk\",\"skill_tools\":\"validate_register.py, evaluate_matrix.py (scientific-brainstorming skill, run in the browser)\",\"register\":{\"absent\":true},\"matrix\":{\"exit\":0,\"weight_delta\":0.1,\"criteria\":[{\"name\":\"information_gain\",\"direction\":\"higher\",\"minimum\":1,\"maximum\":5,\"weight\":3,\"normalized_weight\":0.6},{\"name\":\"resource_burden\",\"direction\":\"lower\",\"minimum\":1,\"maximum\":5,\"weight\":2,\"normalized_weight\":0.4}],\"results\":[{\"idea_id\":\"I1\",\"presentation_rank\":1,\"score\":0.75,\"input_interval\":[0.6,0.85],\"weight_sensitivity_scores\":[0.75,0.75],\"weight_sensitivity_ranks\":[1,1],\"in_shortlist\":true},{\"idea_id\":\"I2\",\"presentation_rank\":2,\"score\":0.65,\"input_interval\":[0.5,0.8],\"weight_sensitivity_scores\":[0.6,0.7],\"weight_sensitivity_ranks\":[2,2],\"in_shortlist\":true}],\"warnings\":[],\"decision\":null},\"shortlist\":[\"I1\",\"I2\"],\"flags\":[{\"id\":\"F1\",\"kind\":\"interval_overlap\",\"text\":\"The input intervals of I1 [0.6, 0.85] and I2 [0.5, 0.8] overlap: the rank order may not hold.\"}],\"browser_status\":\"ready_for_decision_owner\"}",
 "notes": "Imaging core sample-prep session; the decision owner is the core manager."
}

A reply in the lane's contract (a real reply covers every shortlisted idea once):

{
  "task": "review",
  "status": "ready_for_decision_owner",
  "headline": "I1 ranks first on the sheet, but its interval overlaps I2's, so the order is a weak signal for the decision owner.",
  "flag_responses": [
    {"ref": "F1", "stance": "confirmed", "note": "Intervals [0.6, 0.85] and [0.5, 0.8] overlap; the 0.1 score gap rests on one resource_burden point."}
  ],
  "idea_reviews": [
    {"idea_id": "I1", "strongest_version": "Batch staining on a shared rack cuts hands-on time per sample.",
     "counter_observation": "Batching may lengthen the wait for urgent single samples.",
     "alternative_explanations": ["Perceived gain may come from the proposer's own workload.", "Prep time may be driven by booking, not staining."],
     "measurement_failure": "Ratings are board estimates, not timed runs.",
     "sampling_failure": "Three participants from one lab rated it.",
     "harm_or_misuse": "Cross-contamination between samples on a shared rack.",
     "mitigation": "Time ten preps before and after; label rack positions.",
     "residual_uncertainty": "Effect size on turnaround is unknown.",
     "disposition": "retain", "next_action": "pilot-design"},
    {"idea_id": "I2", "strongest_version": "A pre-booked prep calendar removes queueing.",
     "counter_observation": "Unused slots may block other users.",
     "alternative_explanations": ["Queues may already be short outside peak weeks.", "No-shows, not booking, may cause the delays."],
     "measurement_failure": "No queue data was cited.",
     "sampling_failure": "Core users outside the session were not asked.",
     "harm_or_misuse": "None apparent beyond access fairness.",
     "mitigation": "Pull two months of booking logs first.",
     "residual_uncertainty": "Whether queueing is the bottleneck.",
     "disposition": "revise", "next_action": "further-search"}
  ],
  "fragility": [{"idea_id": "I2", "note": "Its weight-sensitivity scores span 0.6 to 0.7; the rank holds at 2 across the tested weights."}],
  "gates": [{"idea_id": "I1", "gate": "biosafety", "status": "unclear", "note": "Depends on which specimens share the rack."}],
  "criteria_issues": ["No criterion covers turnaround for urgent samples."],
  "missing_perspectives": ["Occasional core users", "Biosafety officer"],
  "questions": ["Which specimen classes would share the batch rack?"],
  "decision": null,
  "not_established": ["That either idea reduces prep time.", "That the rank order would survive re-rating by other users."]
}

Worked example: draft

Short notes from a session of three people with three ideas, the ratings written on the board (no register or criteria yet, so none are sent):

{
 "task": "draft",
 "notes": "Session 2026-09-20. Facilitator P1, participants P2 and P3. Question: how can the imaging core cut sample-prep time? Ideas: I1 batch staining on a shared rack (P2); I2 a pre-booked prep-slot calendar (P3); I3 train two more users on the embedder (P1). Criteria: information_gain, higher is better, weight 3; resource_burden, lower is better, weight 2; both rated 1-5. Board ratings: I1 4 and 2; I2 2 and 1; I3 3 and 4."
}

A reply in the lane's contract. The real register is the complete schema 1.1 object the validator reads; here it is cut and marked "...":

{
  "task": "draft",
  "headline": "Three ideas from a three-person imaging-core session, with the board ratings carried into the sheet.",
  "register": {"schema_version": "1.1", "...": "the complete register: the session, pseudonymous participants P1-P3, and ideas I1-I3 exactly as in the notes"},
  "criteria": {"schema_version": "1.0", "criteria": [
    {"name": "information_gain", "description": "How much the idea would teach us", "weight": 3, "direction": "higher", "minimum": 1, "maximum": 5},
    {"name": "resource_burden", "description": "Staff time and cost", "weight": 2, "direction": "lower", "minimum": 1, "maximum": 5}
  ]},
  "scores_csv": "idea_id,information_gain,resource_burden\nI1,4,2\nI2,2,1\nI3,3,4\n",
  "ratings_source": "notes",
  "assumptions_made": [
    {"field": "criteria[0].description", "value": "How much the idea would teach us", "why": "The notes name the criterion but do not describe it."}
  ],
  "open_questions": ["Who is the decision owner for the shortlist?"],
  "changes": []
}

Input fields

Every field is a string - register, criteria and facts included, each a JSON text, not an object; scores is CSV text. This matches the app's declared input schema.

fieldtyperequiredmeaning
taskstringyes"draft" or "review".
notesstringno (draft: yes)Draft: your session notes (the page sends at most 12,000 characters, cut on a sentence and marked [... cut: N more characters not sent]). Review: optional context.
registerstringnoThe idea register as JSON text (schema 1.1 of the scientific-brainstorming skill). Review: when you have one. Draft: only when revising an existing register.
criteriastringno (review: yes)The criteria file as JSON text, schema 1.0: {"schema_version":"1.0","criteria":[{"name","description","weight","direction":"higher|lower","minimum","maximum"}]}. Review: always. Draft: only when revising.
scoresstringno (review: yes)Review only: the scoring sheet as CSV text - the header plus the shortlisted ideas' rows.
factsstringno (review: yes)A JSON-encoded string of what the skill's two tools found - see below.
questionstringnoSomething you want answered inside the lane's fields.
retry_notestringnoOnly when re-asking after a reply that did not parse.

The facts string

facts is a JSON string, not an object. The page computes it by running the skill's validate_register.py and evaluate_matrix.py in the browser. It holds page and skill_tools; register (exit, valid, errors[], warnings[], statistics - or failure, crash or absent); matrix (exit, weight_delta, criteria[] with name, direction, minimum, maximum, weight and normalized_weight, results[] with idea_id, presentation_rank, score, input_interval [lo, hi], weight_sensitivity_scores [min, max], weight_sensitivity_ranks [min, max] and in_shortlist, warnings[], and decision always null); shortlist (idea ids); flags ({id: "F1", kind, text}, ...); and browser_status: ready_for_decision_owner, revise_first or not_reviewable.

If you do not run the tools, you may send your own facts in this shape; the model treats it as given. To compute it outside the page, run the skill's scripts on the exact files you send: python3 scripts/validate_register.py register.json and python3 scripts/evaluate_matrix.py scores.csv --config criteria.json.

The output

A finished run's output text (job.output.output) is one JSON object in the lane's contract, every key present ([] where there is nothing to say). Parse it yourself (step 7), and extract the outermost object defensively.

lanekeycontent
bothtask"draft" or "review": the lane the model answered.
bothheadlineOne sentence.
draftregisterThe complete register object (schema 1.1); participant ids are pseudonymous and every idea comes from the notes.
draftcriteriaThe criteria object (schema 1.0).
draftscores_csvThe scoring sheet as CSV text; blank ratings when the notes give none.
draftratings_sourcenotes | blank.
draftassumptions_made[{field, value, why}]: every value the notes did not state.
draftopen_questionsStrings: what you must decide.
draftchangesStrings: each change against a given register or criteria; [] unless revising.
reviewstatusready_for_decision_owner | revise_first | not_reviewable; never looser than facts.browser_status.
reviewflag_responses[{ref, stance, note}], one per flag in facts.flags; stance is confirmed | dismissed | explained.
reviewidea_reviews[{idea_id, strongest_version, counter_observation, alternative_explanations (at least 2), measurement_failure, sampling_failure, harm_or_misuse, mitigation, residual_uncertainty, disposition, next_action}]; disposition is retain | revise | pause | stop | external-review; next_action is further-search | consultation | simulation | pilot-design | protocol-development | preregistration | no-action.
reviewfragility[{idea_id, note}].
reviewgates[{idea_id, gate, status, note}]; gate is ethics | biosafety | dual-use | human-subjects | animal | data-governance | clinical | regulatory | feasibility; status is review-required | unclear | not-apparent.
reviewcriteria_issues, missing_perspectives, questionsStrings.
reviewdecisionAlways null.
reviewnot_establishedStrings: what nothing here establishes.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"brainstorm-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ....

The input object IS the request body. There is no {"input": ...} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your text.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: register, criteria and facts must be JSON-encoded strings, not objects.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="brainstorm-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://brainstorm-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"brainstorm-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://brainstorm-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"brainstorm-desk"}'
# {"ok":true,"data":{"token":"...","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra. hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": ...} wrapper. /estimate does no input validation, so check the shape yourself: an object whose every value is a string, task equal to draft or review, a draft with non-empty notes (or a register to revise), and a review with criteria, scores and a facts that is a JSON string parsing to an object.

# body.json is the input object itself - no {"input": ...} wrapper (see the
# worked examples above). estimate does not validate it, so check the shape first:
python3 -c '
import json
b = json.load(open("body.json"))
assert isinstance(b, dict) and b.get("task") in ("draft", "review")
assert all(isinstance(v, str) for v in b.values())
if b["task"] == "review":
    assert b.get("criteria", "").strip() and b.get("scores", "").strip()
    assert isinstance(json.loads(b.get("facts", "")), dict)
else:
    assert b.get("notes", "").strip() or b.get("register", "").strip()
'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":...,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, brainstorm-desk:<lane>:<hash>:a<attempt>, so a retried request returns the same job instead of billing a second run. Use one key per distinct input: edited notes, criteria, scores, facts or question are a new hash, and replaying an old key with a different body is a 409. Any stable digest of the body works. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # draft or review
KEY="brainstorm-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"task\":\"review\",\"status\":\"ready_for_decision_owner\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"task\":\"review\",\"status\":\"ready_for_decision_owner\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is one JSON object serialised as a string. Parse it, check that task is the lane you asked for, then use it: for a review, the status, the flag responses and the idea reviews; for a draft, save the register, criteria and scoring sheet as files and run the skill's scripts on them to get the facts for the review.

# The reply is a JSON string inside data.output.output (saved as reply.json in step 5):
python3 -c 'import json;r=json.load(open("reply.json"));print(r["task"],r.get("status",""),r["headline"])'
# A draft: save the three files and compute the facts with the skill's own scripts.
python3 -c '
import json
r = json.load(open("reply.json"))
json.dump(r["register"], open("register.json", "w"), indent=2)
json.dump(r["criteria"], open("criteria.json", "w"), indent=2)
open("scores.csv", "w").write(r["scores_csv"])
'
python3 scripts/validate_register.py register.json
python3 scripts/evaluate_matrix.py scores.csv --config criteria.json

Costs

Invariants worth asserting