# AnalystOS - full documentation > Reads a source document and returns a report in which every figure is a verified quote or an independently recomputed calculation, sealed so anyone can re-verify it offline. Early stage: use non-confidential documents only. # AnalystOS API (v1) Base: `/api/v1`. Machine-readable description: `GET /api/v1/openapi.json`; index of endpoints: `GET /api/v1`. Errors are always `{"error": {"code": "...", "message": "..."}}`. | Method and path | Auth | What | |---|---|---| | `GET /api/v1` | none | index: version, endpoints, limits | | `GET /api/v1/openapi.json` | none | OpenAPI 3.1 | | `POST /api/v1/analyses` | API key | upload one document (`multipart`, field `file`, optional `title`); returns the report, its seal and, if storage is configured, links | | `POST /api/v1/jobs` | API key | same upload; returns immediately, `202 {"id","status":"pending"}` - no model calls yet | | `POST /api/v1/jobs/{id}/run` | API key (creator) | run a pending job (idempotent once no longer pending) | | `GET /api/v1/jobs/{id}` | API key (creator) | poll a job's status; never executes anything | | `GET /api/v1/reports/{id}?format=html\|pdf\|seal` | API key (creator) or signed link | fetch a stored report | | `POST /api/v1/reports/{id}/review` | API key (creator) | `{"approved"?: bool}` -> record that a human reviewed this report | | `POST /api/v1/links` | API key (creator) | `{"id","format","ttl_seconds"}` -> an expiring signed URL | | `GET /api/v1/verify/{id}` | none | seal metadata only (payload, signature info) | | `POST /api/v1/verify` | none | `{"bundle", "public_key"?, "text"?}` -> verification result (stateless) | ## A call, end to end ```bash curl -s -H "Authorization: Bearer $KEY" -F file=@report.pdf https://analystos.dev/api/v1/analyses # -> 201 {"id": "<64 hex>", "tier": "written", "counts": {...}, "seal": {...}, "html": "...", "links": {...}} curl -s -X POST -H "Authorization: Bearer $KEY" -H "Content-Type: application/json" \ -d '{"id":"","format":"html","ttl_seconds":3600}' https://analystos.dev/api/v1/links # -> 201 {"url": "https://analystos.dev/api/v1/reports/?format=html&exp=...&sig=...", "expires": "..."} ``` That `url` is what a FlashyOS task completion takes as `evidenceUrl`; the `id` is the digest to carry alongside it. Whoever receives it can check the seal with `POST /api/v1/verify` or, offline, `python3 -m analystos.l4.seal_verify` (`docs/seal.md`). ## Default: nothing is kept AnalystOS keeps no reports on a server by default (decision by 30E Ventures, 2026-09-26): the upload is deleted when the request ends and the result comes back in the response, for the caller to save. `POST /api/v1/analyses` returns the report (`html`) and its seal (`seal`); add `include_pdf=true` to also get the generated PDF as `pdf_base64`. The hosted upload page does the same and saves the report, the PDF and the seal to the reader's own Downloads folder. For the mesh this means whoever receives a seal hosts it themselves. A seal verifies anywhere (`docs/seal.md`), so a task's `evidenceUrl` can point at the requester's own storage. Server-side storage and signed links, below, are optional and stay off unless `ANALYSTOS_STORE_DIR` is set. ## Async jobs `POST /api/v1/analyses` holds one connection open for the whole pipeline. For an agent that would rather not do that, `/jobs` splits it into three calls: ```bash curl -s -H "Authorization: Bearer $KEY" -F file=@report.pdf https://analystos.dev/api/v1/jobs # -> 202 {"id": "<32 hex>", "status": "pending", "created": "..."} curl -s -X POST -H "Authorization: Bearer $KEY" https://analystos.dev/api/v1/jobs//run # -> 200 {"id", "status": "done", "created", "report_id": "<64 hex>", "links": {...}} curl -s -H "Authorization: Bearer $KEY" https://analystos.dev/api/v1/jobs/ # -> 200 {"id", "status": "done", "report_id", "links"} (or "pending" / "running" / "failed": {"error"}) ``` Once `status` is `done`, `links` and `report_id` work exactly like `/analyses`'s - `GET /reports/{id}`, `POST /reports/{id}/review`, `POST /links`, `GET|POST /verify` - because a finished job's result is stored through the same report store, not a separate mechanism. **This is not a background worker.** `/run` executes the same pipeline `/analyses` does, in one bounded call with the same time ceiling - Vercel runs nothing between requests. What changes is *when* that call happens: the upload is accepted (and can sit safely, briefly, in the job store) before the caller commits to waiting on the actual analysis, and the caller can poll `GET /jobs/{id}` afterward instead of holding the original connection. A document whose analysis alone exceeds the platform's ceiling still fails at `/run` exactly as it would at `/analyses`. Calling `/run` again on a job that is no longer `pending` never re-runs it - it just returns the current status. The uploaded file is kept only between `POST /jobs` and `POST /jobs/{id}/run` (deleted the moment `/run` finishes, success or failure) - the same "nothing kept longer than it has to be" rule as `/analyses`, just delayed by one extra request. A `pending` job that is never run is not purged automatically (see `ROADMAP.md`). ## The `id` `sha256(canonical(seal payload))`. The payload holds the Merkle root, the source and text hashes, the tier, a timestamp and a random nonce, so every analysis gets its own id. ## Configuration (environment) | Variable | Needed for | If unset | |---|---|---| | `ANALYSTOS_API_KEYS` | any authenticated call: comma-separated `name:sha256hex` | every authenticated call is a 503 (fails closed) | | `ANALYSTOS_STORE_DIR` | *optional*: storing reports, links, `GET /reports`, `GET /verify/{id}`, the audit log file | the default: responses carry the report and seal inline; those routes are 503 `store_not_configured`; audit lines go to stderr | | `ANALYSTOS_LINK_SECRET` | `POST /links` and honouring links | 503; no default secret | | `ANALYSTOS_SEAL_KEY` | signing seals (`python3 -m analystos.l4.seal keygen`) | unsigned seals (integrity only), and they say so. A *malformed* value is a 503, not a downgrade. Signed seals can only be authenticated by others once the public key is published (`docs/seal.md`, "The published public key") | | `ANALYSTOS_REPORT_TTL_DAYS` | retention of stored reports | 30 | | `ANALYSTOS_PUBLIC_BASE_URL` | the host used in issued links | derived from the request (`X-Forwarded-*`) | | `ANTHROPIC_API_KEY` | the analysis itself | as for `/api/analyze` | | `ANALYSTOS_MODEL` | *optional*: which model every stage asks (default `claude-sonnet-5`). Recorded in each audit event | the default | | `ANALYSTOS_MAX_IMAGES` | *optional*: the most *distinct* embedded images one PDF may have (each is one model call; a repeated image counts once and an image under half an inch either way is neither counted nor read). More is refused with a clear error, never read in part | 40. A value that is not a whole number from 1 to 1000 is an error | Make a caller key (shown once, stored nowhere): ``` python3 -m analystos.api_v1 newkey partner-name ``` ## Behaviour worth knowing - **Synchronous.** One request runs the whole pipeline and is bounded by the host's function time limit, exactly like `/api/analyze`. `/jobs` (above) splits submit from run from collect, but `/run` is still one such call - see "Async jobs" for exactly what that does and doesn't fix. - **Jobs need `ANALYSTOS_STORE_DIR`.** Unlike `/analyses`, which can run with no store configured (the result just comes back inline), `/jobs` has nowhere to hold the upload between "submit" and "run" without one - `POST /jobs` is 503 `store_not_configured` if it's unset. - **The upload is deleted** when the request ends. Document text is sent to Anthropic's API for the analysis. - **Access.** A stored report is readable by the API key that made it, or via a signed link for one report, one format and an expiry (60 seconds to 7 days, never beyond the report's own expiry). Somebody else's report is indistinguishable from a missing one. Report pages are served with a locked-down `Content-Security-Policy`, `no-store`, and `nosniff`. - **Rate limit.** The same per-IP limiter as the rest of the app (in-memory per process; see `docs/audit-2026-09-25.md` finding 11). - **Storage is not durable on Vercel.** `FileStore` writes to the local filesystem. A production deployment needs a persistent volume, or an adapter with the same methods (`put`, `meta`, `read`, `purge_expired`, `append_audit`, `read_audit`) for an object store. Until then treat stored reports as short-lived. ## Human review The intended shape: an agent calls `/api/v1/analyses` and gets the report back; a human reviews it. `POST /api/v1/reports/{id}/review` (optional body `{"approved": true|false}`) records who reviewed a report and when - same auth and ownership rule as fetching the report. It does not gate anything: a report is deliverable and verifiable whether or not it has been reviewed. The recorded review (or `null`) is returned by `GET /api/v1/verify/{id}` alongside the seal metadata. Reviewing again overwrites who/when/approved. ## Audit log and the charter's measures Every analysis, view, link, review and job appends one JSON line (`audit.jsonl` in the store, else stderr): caller name, report id, tier, fallback reason, proposed / verified / dropped counts, the model used, whether the seal was signed and re-verified. A job run's `analysis` event is identical to `/analyses`'s, plus `"via": "job"`, so it counts toward the measures below the same way; `job_created` and `job_failed` are separate event types. **No document text and no filename.** The three measures in the published charter are computed from it: ``` python3 -m analystos.api_v1 measures [audit.jsonl] ``` Each is `{numerator, denominator, share}` and `share` is `null` on an empty log, never a made-up 100%. A test keeps the measure wording identical to the published charter. ### The log is a hash chain (Slice 84) Each line is chained to the one before it (`seq`, `prev_hash`, `entry_hash`) and, when `ANALYSTOS_SEAL_KEY` is set, signed (`signature`, `key_id`; Ed25519 over the hash with the domain `analystos.audit-entry/1`). With no key, lines are chained and say `"log_signing": "none"`. Check a log offline: ``` python3 -m analystos.api_v1 verify-audit audit.jsonl [--public-key B64URL] [--expect-head HASH] ``` Exit `0` if no check failed, `1` if one did, `2` for unreadable input. Editing, deleting, reordering or inserting a line fails the `chain` check. `authentic` is true only when every line is signed and verifies under a public key **you** supplied. Lines written before the chain existed are tolerated as a prefix; an empty or legacy-only log fails (nothing to verify). Tamper-evident, not tamper-proof: a shortened log is only caught if you pin the `head` you saw earlier (`--expect-head`), and a holder of the signing key can re-sign a rewritten tail. Event fields are unchanged; the chain fields are added. ## Vercel routing Vercel maps one file to one path, so each route has a door in `api/v1/` (`index`, `analyses`, `jobs`, `reports`, `links`, `verify`, `openapi`) that imports the shared Flask app, and `vercel.json` rewrites `/api/v1/openapi.json`, `/api/v1/reports/:digest`, `/api/v1/reports/:digest/review`, `/api/v1/verify/:digest`, `/api/v1/jobs/:job` and `/api/v1/jobs/:job/run` onto them. The routes accept both the pretty path and the `?digest=`/`?id=` form. **Verified on the deployed site, 2026-09-26:** `/api/v1`, `/api/v1/openapi.json` and both forms of the `reports` and `verify` routes reach the app (an unconfigured store answers 503 `store_not_configured`, an unconfigured key set answers 503 `unavailable`, as designed). Not yet exercised in production: a real analysis upload, and the jobs routes added in Slice 74. --- # The AnalystOS seal A seal lets someone who did not run AnalystOS check a delivered report without trusting AnalystOS. It is a JSON bundle; checking it needs only `analystos/l4/seal_verify.py` (Python standard library, plus `cryptography` for the signature) or your own port of the ten rules below. ``` python3 -m analystos.l4.seal_verify bundle.json \ [--public-key ] \ [--text extracted.txt] ``` Exit `0` if no check failed, `1` if one did, `2` for unreadable input. Output is JSON: `{"ok", "authentic", "content_checked", "checks": [{"name","status","detail"}]}`, each status `pass`, `fail` or `skipped`. ## What each level means | Flag | True when | It does **not** mean | |---|---|---| | `ok` | No check failed: the bundle is internally consistent. | The report is right, or that AnalystOS made it. | | `authentic` | `ok`, and the Ed25519 signature verifies under a public key **you supplied**. | A key that came inside the bundle proves nothing: anyone can sign with their own key and include it. | | `content_checked` | `ok`, and with the extracted text: it is the sealed text, every citation is in it, every number equals a number its citation spells, every calculation recomputes. | That the *choice* of figures is good. | The `checks` list, in order (each `pass`, `fail` or `skipped`): | Check | Needs the text? | What it proves | |---|---|---| | `structure` | no | it is an `analystos-seal/1` bundle with every required payload field (and any accountability field present is a string or null) | | `fact_hashes` | no | every fact hashes to its recorded hash; no duplicate keys | | `merkle_root` | no | the facts make the signed root; `entries` is right | | `report_hash` | no | the report is the one that was sealed (or there is none) | | `report_refs` | no | every fact the report cites was sealed | | `signature` | no | Ed25519 over the payload; `pass` only under a key you pinned | | `source_text_hash` | yes | the supplied text is the sealed text | | `citations_in_text` | yes | every citation is in the text | | `values_match_citations` | yes | every quote value and calculation operand equals a number its citation spells | | `calculations` | yes | every calculation recomputes from its operands | An unsigned bundle (no `ANALYSTOS_SEAL_KEY` when it was made) is `ok` but never `authentic`. AnalystOS never signs with a built-in default key. ## Bundle format (`analystos-seal/1`) ```json { "format": "analystos-seal/1", "payload": { "version": 1, "org": "analystos", "entries": 2, "created": "2026-09-25T00:00:00Z", "nonce": "<16 hex chars, random per seal>", "root": "", "source_sha256": "", "text_sha256": "", "report_sha256": " or null", "tier": "written | deterministic | plain", "model_id": " or null", "caller_id": " or null", "code_version": " or null" }, "facts": [ { "key": "fact/000000", "record": { ... }, "hash": "" } ], "report": { ... } or null, "signature": { "alg": "ed25519", "sig": "", "public_key": "", "key_id": "<16 hex>" } or null } ``` `nonce` makes every seal's payload (and therefore the report id, `sha256(canonical(payload))`) unique even for identical content in the same second. `tier` is inside the signed payload so a fallback report cannot be relabelled as a written one. `source_sha256` is the same SHA-256 the evidence store names files by. `model_id`, `caller_id` and `code_version` (Slice 83) say who and what produced the report, inside the signed payload so a signed seal cannot be relabelled after the fact: - `model_id` is the model id the run asked for (`ANALYSTOS_MODEL`, read when the run finished; the same value the audit log records). It is the *requested* id, not a statement from the model provider. - `caller_id` is the `/api/v1` API-key **name** that made the run (an operator-chosen label, never the key). Runs outside `/api/v1` carry `"(cli)"` or `"(legacy-access-code)"`; those contain parentheses, which no key name can, so they never collide with a real caller. `null` means the sealer did not say. - `code_version` is the deployed code's commit id or label, taken from the deployment's environment (`ANALYSTOS_CODE_VERSION`, else Vercel's `VERCEL_GIT_COMMIT_SHA`), never looked up at request time. `null` means the deployment does not say. **Compatibility.** These three fields are an additive change to `analystos-seal/1`, with no version bump. They are optional: a seal made before Slice 83 lacks them and still verifies, and a verifier that ignores unknown payload fields (as `seal_verify.py` always did) verifies a new seal unchanged, because the signature covers the whole payload including them. A *strict* port that rejected unknown payload keys would reject new seals; ports should ignore unknown payload fields, and may check that each of these three, when present, is a non-empty string (at most 128 characters) or `null`. As with every field here, a field is only as trustworthy as the key that signed it: unsigned seals carry them unauthenticated. ## Fact records `facts[].record` is one of four shapes, discriminated by `type`. Numbers are already strings (rule 1). The machine-readable definition is the `SealFact` schema in `https://analystos.dev/api/v1/openapi.json` (`components.schemas.SealFact`, JSON Schema 2020-12 keywords); validate a bundle's facts against it before running the ten rules. It checks **shape only**: a record can fit it and still fail every content check. | `type` | Required | Also present | |---|---|---| | `quote` | `citation` | `value` (numeric string, must equal a number `citation` spells), `sentence` with one `{value}`, `text` when there is no value, `label`, `format`, `display` | | `computed` | `operation`, `value`, `operands`, `total`, `citation` (array) | `sentence`, `label`, `format`, `display`. `operation` is one of `sum average ratio growth_percent percent_of_total difference remainder`; `citation[i]` is operand `i`; `total` is null except for `percent_of_total` | | `event` | `what`, `citation` (equals `what`) | `date`, `status`, `next_step`, `milestones` (`{date, detail}` list); each part is a substring of the document | | `prose` | `text` | none of its own; it carries no figure | Every type may also carry `horizon` (`reported`, `guidance`, `projected`) and `gaap_status` (`gaap`, `non_gaap`, `n/a`). "Required" is what a verifier reads; unknown extra properties are allowed and ignored. ## The ten rules 1. **Normalize.** In every `record` and in `report`, each JSON number is written as a string (`repr` of a float, `str` of an int). Booleans, null, strings, arrays and objects stay. This is why no verifier has to agree with another about how to print a float. **A verifier rejects a bundle** (a `structure` failure) in which any `record`, the `report`, or any `payload` field other than `version` and `entries` contains a JSON number. Those two are the only numbers in a bundle and must be integers: `1`, not `1.0`. 2. **Canonical JSON.** Keys sorted at every level, separators `,` and `:` with no whitespace, UTF-8, non-ASCII characters not escaped. 3. **Fact hash.** `sha256(canonical(record))`, lowercase hex. 4. **Leaf.** `sha256(0x00 || utf8(key + "\0" + fact_hash))`. 5. **Node.** `sha256(0x01 || left_bytes || right_bytes)` over the 32 raw bytes. 6. **Root.** Sort leaves by `key`; pair neighbours left to right; an odd node at the end of a level is promoted unchanged; repeat until one is left. `entries` is the number of facts. (This is the construction in Flashy's own published provenance code, `flashyos-wdk` `provenance.ts`.) 7. **Report hash.** `sha256(canonical(report))`, or `null` with no report (the plain tier). A report that cites `{{N}}`, or `fact_index` / `value_fact` / `delta_fact` N, must cite a sealed fact (`N < entries`). 8. **Signature.** Ed25519 over `canonical(payload)`, base64url without padding. 9. **Text.** `text_sha256` must equal `sha256(utf8(extracted_text))`. 10. **Content.** With the text, fold both text and each citation (lowercase; `–`, `—`, `−` to `-`; drop `$` and `|`; remove a thousands comma between digits; collapse whitespace). A citation passes if the folded citation, or the folded citation with one trailing scale word (`thousand million billion bn mm k m b`) removed, occurs in the folded text **as a whole number**: if it starts with a digit, no ASCII digit, and no `.` or `,` that follows a digit, may come immediately before it; if it ends with a digit, no digit, and no `.` or `,` that is followed by a digit, may come immediately after it. A full stop or comma that merely ends a sentence or separates a list does not count. If one occurrence fails this test, later occurrences are tried. **Values:** read every number a citation spells (optional `(` or `-` makes it negative; optional `$`; digits with commas and a decimal point; an optional scale word `k thousand m mm million b bn billion` multiplies it). The document's *declared scale* is the largest of "in thousands/millions/billions" or "thousands/millions/billions of [U.S.] dollars" found in the text, else 1. A quote's `value`, and each calculation operand's value (`citation[i]` is operand `i`; a `percent_of_total` total is the citation after the operands), must equal one of those numbers as printed or times the declared scale. A `computed` fact then passes if its operation recomputes from its `operands`: `sum`; `difference` and `remainder` = first minus the rest; `average`; `ratio` = first / second; `growth_percent` = (second - first) / first x 100 (either operand order); `percent_of_total` = operand / total x 100 - within relative 1e-9. ### Worked vector Two facts, keys `fact/000000` and `fact/000001`: | key | canonical record | fact hash | |---|---|---| | `fact/000000` | `{"text":"Momentum continued.","type":"prose"}` | `86c4578bbaddfca5459cf5c8f86be65c3eb6aa12153a57749b74f8c5180a7c0b` | | `fact/000001` | `{"citation":"$498.0 million","label":"Revenue","type":"quote","value":"498000000.0"}` | `94e5d1320d014c7eaf7e500190415425449d075fd49b37b7e367543685b23d3a` | Leaves: `788741d444405b11399adb37f76b01291428567136aec8909efac24a210dfe3a` and `289899dacf126d005f2d43c47230b956fc51549a4910d369fbe01853c89f5970`. Root: `7093b80cc076808b59c1500f0cc425003408f757add4afad1963e40c8e044747`. If your port does not reproduce that root, it is not the same seal. ## Making one - The CLI (`python3 -m analystos `) writes `section.seal.json` next to `section.html` for every narrated run. - From Python: `build_report(path, ..., trace=t)` then `analystos.l4.seal.build_bundle(t, signing_key=...)`. - Signing: `python3 -m analystos.l4.seal keygen` prints a new key pair once and stores nothing. Put the secret line in `ANALYSTOS_SEAL_KEY` and publish the public half as described in "The published public key" below. A malformed key is an error, not a silent downgrade to unsigned. ## The published public key A signature only counts under a key you obtained **independently of the seal** (a key inside the bundle proves nothing). When signing is on, AnalystOS publishes its public key at https://analystos.dev/.well-known/analystos-seal-key.json and **until that file exists, seals are unsigned and cannot be authenticated by anyone**. (The file is added by an operator, not by the code; see "Turning it on".) ```json { "format": "analystos-seal-key/1", "keys": [ {"added": "2026-10-02", "algorithm": "ed25519", "key_id": "<16 hex>", "public_key": "", "status": "active"} ], "org": "analystos", "seal_spec": "https://analystos.dev/docs/seal.md" } ``` Rules a reader (or a port) should enforce: `format` is exactly `analystos-seal-key/1`; `public_key` is the canonical unpadded base64url of 32 bytes; `key_id` is the first 16 hex characters of `sha256(those 32 bytes)` and equals the `key_id` in a seal's `signature` block; ids are unique; at least one key is `active`; the file never contains a private key (`validate_key_file` rejects fields named `private_key`, `seed`, `secret` and similar). Unknown extra fields are ignored. Verify a seal against it: ``` curl -sO https://analystos.dev/.well-known/analystos-seal-key.json python3 -m analystos.l4.seal_verify section.seal.json --key-file analystos-seal-key.json [--text extracted.txt] ``` The key is selected by the seal's `key_id`. A seal whose key is **not in the file**, a seal from another `org`, and an invalid file are all **failures** (`key_file` check), never a silent check without a pinned key. An unsigned seal cannot be authentic. **What this does and does not give you.** The file is trusted exactly as far as a TLS connection to analystos.dev is: whoever controls the site can publish a different key. For more than that, record the `key_id` out of band when you first adopt AnalystOS (a contract, a config in your own repository) and compare it before trusting a file. **Rotation** adds a new `active` key and marks the old one `retired` (old seals keep verifying). **Revocation is removal**: delete a compromised key from the file and its seals stop being authentic. ### Turning it on (an operator's steps) Nothing in the repository does this for you, and nothing here ever needs a private key in a file, a chat or git. 1. In your own terminal: `python3 -m analystos.l4.seal keygen`. It prints the **private** seed (`ANALYSTOS_SEAL_KEY=...`) once and stores nothing. Put it straight into your host's secret store (Vercel: Project Settings, Environment Variables, Production, `ANALYSTOS_SEAL_KEY`, sensitive) and a password manager. Never commit it. 2. In the same terminal, with the seed in the environment, write the **public** file (the seed is never printed or written): `ANALYSTOS_SEAL_KEY= python3 tools/seal_key_file.py --from-env` (or `--public-key `; the tool refuses a value equal to the seed in your environment). It creates `site/.well-known/analystos-seal-key.json`. Then close the terminal. 3. `python3 tools/refresh.py`, run the tests, commit the file (public data only), open the pull request and merge. Redeploy so the new environment variable is live. 4. Prove the two halves match: run one analysis in production, download its seal, and run `python3 tools/smoke.py --expect-seal-key --seal ~/Downloads/.seal.json`. It must report `a seal from this deployment verifies under the published key`. This is the check for the one dangerous mistake, a signing key in the host that is not the key on the site. 5. To rotate later: generate a new key, run `tools/seal_key_file.py --from-env --add` (new key active, old one retired), deploy the file, then switch `ANALYSTOS_SEAL_KEY`. `tools/seal_key_file.py --check ` validates a file. Either half alone is safe: the key in the environment without the file means signed seals nobody can yet authenticate; the file without the key means unsigned seals. ## Limits, stated plainly - The extracted text is **not** in the bundle (it may be confidential). To check content you must supply text whose hash matches; otherwise the hash check fails and the content checks still run on what you supplied. - Event facts (dates, milestones) and prose carry no numeric value; the checks above do not cover them beyond the citation being in the text. - The seal proves the report was built from these facts. It does not prove the facts are the important ones, that a relationship between two figures is right (Slice 55 checks some), or that the model's GAAP/guidance labels are correct. - There is no chain between reports and no transparency log yet. - A citation with its scale word stripped (rule 10) matches any standalone occurrence of the bare number, even one the scale was never meant to apply to - `$5 million` is satisfied by a document that only ever says "5" on its own. Rule 10 proves the number is genuinely present as a whole number, not that it is the right one; that is still `values_match_citations`'s job where it applies, and a human's job otherwise. --- # AAO charter checking `analystos/aao/` checks a FlashyOS AAO 0.1 charter without Node, npm or the network. It exists so AnalystOS can prove its own charter is well-formed before asking anyone else to read it, and so a second implementation exists to compare with FlashyOS's own. ## Use ``` python3 -m analystos.aao path/to/charter.json # or "-" to read stdin ``` Prints one JSON object; exit code `0` valid (warnings allowed), `1` invalid, `2` unreadable input. ```json { "valid": false, "errors": [{ "code": "aao.schema.enum", "level": "error", "path": "/roles/0/humanApprovalAtOrAbove", "message": "must be one of ['LOW', 'MEDIUM', 'HIGH', 'CRITICAL']; got 'NONE'" }], "warnings": [], "schema": { "source": "https://flashyos.com/aao.schema.json", "retrieved": "2026-09-25", "sha256": "99f520bb..." } } ``` From Python: ```python from analystos.aao.validate import validate_charter, errors, is_valid problems = validate_charter(doc) # list of Problem(code, level, path, message) is_valid(doc) # True when there is no *error* ``` `validate_manifest(doc)` is the older string-list form (errors only). ## With FlashyOS's conformance-kit ``` npx @flashyos/conformance-kit -- python3 -m analystos.aao.kit ``` The adapter follows the kit's line protocol (`{id,set,input}` in, `{id,valid,codes}` out). A document with only warnings is valid. **No AAO corpus is public**: the kit ships one bundled corpus, for `frontdoor/1`, which this adapter would (correctly) disagree with everywhere - it checks charters. ## What is checked | Layer | What | Source | |---|---|---| | Schema | types, `aao` = `"0.1"`, slug/role-name patterns, role-name length 3-24, the ten closed families, impact tiers LOW/MEDIUM/HIGH/CRITICAL, email shape of `accountableTo`, `x-` extensions allowed, any other unknown key refused | `analystos/aao/aao.schema.json`, a pinned byte copy (sha256 in `validate.py`) of the published schema | | Cross-field | duplicate role names; `worksIn` naming an undeclared repository; `escalation` naming a missing role; more than one default repository; a repository no role owns | the schema's own description text | | Naming | at most three words; a **partial** vendor list; placeholder names; `you@example.com` | FlashyOS's documented role-naming rules and the `accountableTo` description | | Advice | a role with no `measure` (warning) | schema: "strongly encouraged" | Not checked: FlashyOS's codename denylist (we do not have it - a role called `nova` passes here and may not pass theirs), and whether an org's handshake, front door or live behaviour conform (that is what `npx @flashyos/conformance` grades; see `docs/flashyos-alignment-2026-09-25.md`). ## Rule codes These are AnalystOS's, not FlashyOS's (their corpus is not public, so their codes are unknown to us). | Code | Level | Meaning | |---|---|---| | `aao.schema.type` | error | wrong JSON type (or the document is not an object) | | `aao.schema.const` | error | `aao` is not `"0.1"` | | `aao.schema.enum` | error | not one of the allowed values (family, tier) | | `aao.schema.pattern` | error | string does not match its pattern (slug, role name, email, blank text) | | `aao.schema.length` | error | role name shorter than 3 or longer than 24 | | `aao.schema.min_items` | error | an array that must not be empty is empty | | `aao.schema.required` | error | a required key is missing | | `aao.schema.unknown_key` | error | a key that is neither defined nor `x-` prefixed | | `aao.role.duplicate_name` | error | two roles share a name | | `aao.role.worksin_undeclared` | error | `worksIn` names an undeclared repository | | `aao.role.name.too_many_words` | error | more than three hyphen-separated words | | `aao.role.name.vendor` | error | a role name uses a model-vendor word | | `aao.role.name.placeholder` | error | a role name is a placeholder | | `aao.escalation.unknown_role` | error | `escalation` names no role | | `aao.repository.multiple_default` | error | more than one default repository | | `aao.repository.no_owner` | error | a repository no role lists in `worksIn` | | `aao.accountable.placeholder` | error | `accountableTo` is `you@example.com` | | `aao.role.no_measure` | warning | a role declares no `measure` | ## Updating the pinned schema `aao.schema.json` must stay byte-identical to the published file. To update it, fetch the new one, replace the file, then change `SCHEMA_SHA256` and `SCHEMA_RETRIEVED` in `validate.py` in the same commit. `tests/test_aao_validate.py` fails if the hash drifts, and if the new schema uses a keyword the checker does not implement. --- # Mesh identity files What AnalystOS publishes so a FlashyOS conformance check (or any agent) can discover it, and what it deliberately does not. | URL (once deployed at analystos.dev) | File in this repo | Format | Purpose | |---|---|---|---| | `/.well-known/flashyos.json` | `site/.well-known/flashyos.json` | `flashyos/1` handshake | Level 1: discoverable | | `/.well-known/flashyos-charter.json` | `site/.well-known/flashyos-charter.json` | AAO 0.1 charter | Level 2: chartered | | `/flashyos.roles.json` | `site/flashyos.roles.json` | same charter, legacy path | some estates serve both; a test keeps the bytes identical | | `/.well-known/api-catalog` | `site/.well-known/api-catalog` | RFC 9727 linkset | points agents at the OpenAPI document and the API guide (generated by `tools/build_site_machine.py`) | `vercel.json` serves these with the right content type (JSON, or `application/linkset+json` for the catalog), `Access-Control-Allow-Origin: *` and a 5-minute cache. Other machine-facing files (`/llms.txt`, `/llms-full.txt`, `/docs/*.md`, `/sitemap.xml`, `/robots.txt`) are generated from `docs/` by `tools/build_site_machine.py`. ## The charter in one screen Slug `analystos`; accountable human `30eventures@gmail.com`; escalation goes to `verification`. Three standing roles, each with one number it moves. Every measure must be computable from what the API's audit log records (`docs/api.md`), so nothing is published that cannot be measured. | Role | Family | Approval at or above | Measure | |---|---|---|---| | `analysis` | data | MEDIUM | reports delivered at the written tier, both gates passed, as a share of reports delivered | | `verification` | risk | HIGH | facts refused for failing verification, as a share of facts proposed | | `evidence` | governance | HIGH | delivered reports whose seal re-verifies offline, as a share of reports delivered | ## The handshake ```json { "mesh": "flashyos/1", "org": { "slug": "analystos", "name": "AnalystOS", "profile": "https://analystos.dev" } } ``` No `capabilities`: a capability is a claim that something is callable. Add `"capabilities": ["document-analysis"]` only after `docs/api.md`'s endpoints are deployed and an agent token exists. ## Check it ``` python3 -m analystos.aao site/.well-known/flashyos-charter.json # ours (docs/aao.md) npx @flashyos/conformance analystos.dev --level 2 # FlashyOS's ``` To check a whole deployment (every advertised URL and content type, the charter, the API's rewrites and that it fails closed), run one command: ``` python3 tools/smoke.py https://analystos.dev # add --json, or --no-post ``` If Python reports a certificate error, the machine's Python has no CA bundle (the python.org macOS build): use `SSL_CERT_FILE=/etc/ssl/cert.pem python3 tools/smoke.py`. Run against analystos.dev on 2026-09-27 it passed 54 of 54 checks. **Result, 2026-09-26:** the second command was run against the deployed site (`@flashyos/conformance` 0.2.3, run with install scripts disabled) and exited 0: Level 1 (Discoverable) and Level 2 (Chartered) both passed, every check ticked. The package's published code was read before it was run: it makes only GET requests to the domain (and, at Level 3, to `api.flashyos.com`'s public conformance record), reads no local files for the check, and has no install scripts. The source repository is private, so the published code was read, not its source. Level 3 (the mark) is granted from FlashyOS's register and needs a running agent; it is not attempted here. To repeat the run: ``` npm_config_ignore_scripts=true npx @flashyos/conformance@0.2.3 analystos.dev --level 2 ``` ## Deliberately not served | File | Why not | |---|---| | `/.well-known/frontdoor.json` | `frontdoor/1` requires a working `endpoint` and a person who reads what arrives. Neither exists. | | `/directory.fragment.json` | asserts people and relationships on someone's authority; the owner's decision. | | `/.well-known/canon.json`, `backlog.json` | nothing to pin or publish yet. | ## Ownership and contact (stated by the owner, 2026-09-26) 30E Ventures owns AnalystOS and the `analystos` org on the FlashyOS network, and `30eventures@gmail.com` is the accountable email. This is the owner's statement; the public directory does not list org owners, so it is confirmed for real when someone signs in at app.flashyos.com as that org. ## Open decisions before deploying 1. ~~Who owns the existing `analystos` org?~~ Answered above: 30E Ventures. Still worth confirming by signing in as that org before the handshake goes live. 2. ~~Is `30eventures@gmail.com` the right accountable human?~~ Yes, per the owner. 3. **Are the measures right?** They are computable, not yet computed. ## Not verified Verified against the deployed site on 2026-09-26: Vercel serves `.well-known` and the headers apply (JSON, `application/linkset+json`, CORS); our checker and FlashyOS's Level 2 check both pass the live charter. Still not verified: - A real upload through `/api/v1/analyses` or the upload page's new download buttons (needs an access code or API key, and spends real model calls). - That the `analystos` org is controlled by 30E Ventures on the FlashyOS network (the owner's statement; confirm by signing in at app.flashyos.com). --- # Architecture AnalystOS is a layered stack. Each layer adds one thing; the layer below it is its source of truth. | Layer | Name | What it does | |-------|------|--------------| | L0 | Evidence kernel | Every source stored once, addressed by a hash of its content. Immutable. | | L1 | Structure layer | Turns documents and tables into typed, validated data. Corrections train it. | | L2 | Reasoning layer | Retrieval + analysis over the evidence graph. Every claim points back into L0. | | L3 | Workspace | Where the analyst works: canvas, model, memo, review. Reads like a document. | | L4 | Deliverable & attestation | Exports working papers etc. with a verifiable evidence trail attached. | | L5 | Distribution | Multi-tenant control plane: KPMG, Gord, public. Per-tenant isolation. | | L6 | AAO manifest & mesh | AnalystOS as an actor on the FlashyOS mesh, callable by other agents. | Running through every layer: identity & access, an append-only audit log, a policy engine (which data may reach which model), and observability. ## Frozen bits - **The L0 provenance model.** Once analysis code depends on how sources are hashed and referenced, that scheme does not change. Flag problems; do not redesign it. ## Current state What exists, by layer (code in `analystos/`; roadmap in `ROADMAP.md`). "Built" means implemented and tested here; it says nothing about production. | Layer | Status | Where | |---|---|---| | L0 Evidence kernel | Built: SHA-256 content addressing, Fernet encryption at rest (key you set; a public development fallback if you do not), retention purge | `analystos/l0/` | | L1 Structure | Built: CSV, Excel, Word, PowerPoint, PDF (multi-column, image-derived facts), table unification, footnotes, boilerplate stripping | `analystos/l1/` | | L2 Reasoning | Built: a model proposes facts, code verifies them (quote, computed, event, prose); a narrator writes from verified facts; Gate 1 validators; a blind Gate 2 proofreader | `analystos/l2/` | | L3 Workspace | Not built | - | | L4 Deliverable and attestation | Built: HTML and PDF reports, charts, a zero-model deterministic floor, basis tags, **sealed reports** with a standalone verifier | `analystos/l4/` | | L5 Distribution | Not built: one shared access code plus per-caller API keys; no tenants | `api/`, `analystos/api_v1/` | | L6 AAO and mesh | Partly built: an AAO checker, a `flashyos/1` handshake and charter (in the repo, not deployed), an agent-callable API. No org, token or capability on the network | `analystos/aao/`, `site/.well-known/`, `analystos/api_v1/` | Cross-cutting: an append-only audit log (API only), a per-IP rate limiter (in-memory), an in-code API-call budget. Not built: identity and access beyond API keys, a policy engine, observability. Where to read next: [`api.md`](api.md) (calling it), [`seal.md`](seal.md) (checking a report), [`aao.md`](aao.md) and [`mesh-identity.md`](mesh-identity.md) (the FlashyOS side), `docs/decisions.md` (in the repository) (why), and the code-verified `docs/audit-2026-09-25.md` (in the repository). --- # Using AnalystOS AnalystOS takes a table and a list of questions, and writes a short section of prose where **every number has a footnote back to the exact cell it came from**. Anyone can check your figures without asking you. Everything is a command you type in the **Terminal** app. Three steps. --- ## Before you start - Your data must be a **CSV file**. In Excel: *File → Save As → CSV*. - Open Terminal and go to the project folder — **do this first, every time**: ``` cd ~/AnalystOS/AnalystOS ``` If a later command says *"No such file or directory"*, you're probably in the wrong place — run that line again. --- ## Step 1 — make a job from your CSV ``` python3 -m analystos.scaffold path/to/your.csv jobs/my-first-job ``` This creates a folder `jobs/my-first-job/` holding a copy of your CSV and a `job.json` file. It prints the next command to run. --- ## Step 2 — edit the job Open `jobs/my-first-job/job.json` in a text editor (TextEdit, VS Code — **not** the Terminal). There are two ways to fill it in. ### The fast way: a template If your table looks like an income statement (a period column, and any of revenue / cost of revenue / gross profit / operating expenses / operating income / net income, however they're labeled), skip `asks` entirely: ``` { "title": "Review of your.csv", "source": "your.csv", "schema": { "period": "text", "revenue": "number", "net_income": "number" }, "template": "income_statement", "currency_unit": "actual" } ``` This generates the standard questions automatically — the latest period's figures, year-over-year growth, margins — for whichever line items it recognizes. Nothing is invented: a line item it doesn't recognize is simply left out. **`currency_unit` matters and only has one right answer per table.** It says what scale the numbers in your table *are* — look at what your source document itself says (a real income statement almost always states this, e.g. "$ in millions"): - `"actual"` — the numbers are already raw dollars (`4200000` means $4.2M). - `"thousands"` — the numbers are in thousands (`4200` means $4.2M). - `"millions"` — the numbers are in millions (`4.2` means $4.2M). Get this wrong and every dollar figure in the report is confidently wrong by a factor of 1,000 or 1,000,000 - not an error, just a wrong number. Check it against the source before you run the job. ### The manual way: write your own questions For anything a template doesn't cover, write `asks` yourself: ``` { "title": "Review of your.csv", "source": "your.csv", "schema": { "period": "text", "revenue": "number" }, "currency_unit": "actual", "asks": [ { "text": "period Q1 revenue was {answer}.", "format": "usd", "where": ["period", "Q1"], "select": "revenue" } ] } ``` - **`title`** — what this section is about. - **`asks`** — one entry per fact you want. Each entry has: - `text` — the sentence to write. `{answer}` is where the number lands. - `where` — `[column, value]`: find the row where *column* equals *value*. - `select` — the column whose value is the answer. - `format` (optional) — see below. **Worked example** — *"give me revenue for the FY2024 row"*: ``` { "text": "FY2024 revenue was {answer}.", "format": "usd", "where": ["period", "FY2024"], "select": "revenue" } ``` Add as many `asks` as you want. Put a comma between entries; the **last** entry has no comma after it. **Making numbers readable.** Add an optional `"format"` to any ask so the number prints properly instead of raw (`4200000.0`): - `"usd"` — a dollar amount: prints as `$4.2M` / `$1.3B`, scaled by the job's one `currency_unit` (see above) — never set per-ask. - `"percent"` — adds a `%` sign (growth and margin asks return a percent number). - `"number"` — adds thousands commas: `4,200,000`. A negative value always prints in parentheses — `($4.2M)` — the standard way of showing a loss. --- ## Step 3 — run it ``` python3 -m analystos jobs/my-first-job ``` It prints the section and writes two files into `jobs/my-first-job/`: - `section.md` — the plain-text version. - `section.html` — a formatted page, which **opens in your browser automatically**. To get a PDF: in the browser, **File → Print → Save as PDF** (or press **Cmd+P**). --- ## Reading the output ``` FY2024 revenue was 4200000.0. [1] --- [1] source 32110ff4... - row 4, column "revenue" ``` The `[1]` footnote means: this number came from **row 4** (the header counts as row 1), **column "revenue"**, of the file whose contents fingerprint to `32110ff4...`. That fingerprint changes if even one character of the source changes — so a footnote that still matches is proof the source wasn't altered. --- ## If something goes wrong - **Pasted JSON into the Terminal and it hung** (a `>` or `cursh>` prompt) — press **Ctrl+C**. `job.json` is a *file*; edit it in a text editor. - **"No such file or directory"** — run `cd ~/AnalystOS/AnalystOS` first. - **`no row where 'period' == 'FY2024'`** — the value in your `where` doesn't match the CSV exactly. Check spelling, spaces, and things like `FY24` vs `FY2024`. --- ## What it can't do yet - Only **CSV** files. - Only **exact-match lookups** — no sums, growth rates, or ratios. - No plain-English questions — you fill in `where` and `select` yourself. --- ## Reading a document instead of a table Running `python3 -m analystos ` on a job with no `asks` and no `template` reads the document's real text (CSV, Excel, Word, PowerPoint or PDF), proposes facts with a model, verifies each in code, and writes a report. It needs `ANTHROPIC_API_KEY` in your environment. Besides `section.html` and `section.pdf` it writes `section.seal.json`: a seal you or anyone else can check with `python3 -m analystos.l4.seal_verify section.seal.json` (see [`seal.md`](seal.md)). Set `ANALYSTOS_SEAL_KEY` first to sign it. To call AnalystOS from another program instead, see [`api.md`](api.md). ---