ReLoop.me · Profile Intel & Role Intel

Methodology

What the classification vocabularies are, what the system checks before it delivers, what it confirms versus what it infers, and what we have actually measured about whether it repeats itself. Including the claims we have withdrawn.

Hand-maintained. Not generated from the pipeline. Updated weekly — the positioning overview at reloop.me/methodology covers the two-sided model and updates on a slower, monthly cycle; this page is the technical audit underneath it. Last updated 2 September 2026. Profile Intel engine on this page: prompt 3.27.2 · Quality Reviewer 3.20.12 · Salary 2.49 · Object Transformer 1.14.6 · Renderer 16.64. Role Intel engine on this page: renderer 27.9.58 · delivery gate 1.8 · salary 3.50 · object transformer 2.4.3 · request builder 4.10.0 · reviewer prompt 2.9 · input calibrator prompt 1.1. All eight read back from the live workflow on 2 September 2026, not from notes.

The two products

Both are read from a document you submit, both are delivered to this domain, and both run on the same engine. What differs is what they read, what they emit, and how much of each has been measured. They are set side by side so the difference is visible rather than asserted — including where one column is fuller than the other.

There is a second reason they sit together. Both already emit the same class of structured output — a closed classification, an evidence trail, a pay figure with its basis stated — and that shared shape is what a future comparison layer would read. Role × Profile, the planned layer that reads a person's structured profile directly against a role's structured demand, does not exist yet. When it does, it will not re-read either document; it will read the two structured outputs this page already documents. Nothing below should be taken as evidence that layer is built — the open backlog on each side, and the "On the roadmap" section further down, are what is actually true today.

Status at a glance

ProductComparable todayOpen backlog
Role Intel4 of 7 fields — the money and what feeds it, not the archetype word10 items
Profile Intel7 of 9 fields — the money and the vocabulary, not management tier or peak reports5 items

Counted by hand from the comparability tables and open-backlog lists below on every weekly update — not computed, and not a substitute for reading either list.

Role Intel €7 Measured 2 Sep 2026

A structured read of a role — what a job description is actually asking for, and how to read it.

Closed vocabularies

Same test as the column beside this one: for each value list, does the prompt lock it or merely list it. Role Intel answers worse than Profile Intel on the first vocabulary, and that is the most important line on this page.

Role archetypes 7, listed — not locked

Builder · Operator · Specialist · Connector · Fixer · Translator · Anchor

The prompt carries the seven bare words and no definitions. The definitions exist only as tooltip strings in the renderer, written after the fact, and the model never sees them. There is no runtime check that the emitted value is one of the seven. A role carries a dominant archetype and may carry a secondary.

Seniority ladder 5 producible — not validated

Senior IC · Manager · Senior Manager · Director · VP/Executive

These five are the rungs the pay engine can actually price and label. Three problems sit behind that sentence. The free-text matcher accepts seventeen tokens and maps them onto those five. ic_junior is priced but has no band label, so a junior role is captioned with the senior band's name, and it is absent from the ladder, so it can never be scope-adjusted. And ic_mid, c_level and head_of are referenced by three consumers but producible by none.

Occupational family · geography closed

Families resolve through an alias map to a canonical set. Geography resolves to one of ten bands: four Romanian cities plus RO_other, and EU_west, EU_cee, EU_south, EU_nordic_ch, EU_unspecified. Pay is held per family × geography × rung.

Pay confidence closed

Calibrated or Estimated, with Disclosed where the employer published a band. Any inferred seniority now forces Estimated — that was shipped on 1 September 2026 and was not true before it.

Two vocabularies are not closed, and one of them is not a vocabulary at all. gap_type — the label naming what a role is missing — has no fixed list: two runs of the same input produced eight distinct labels with one overlap. And role_seniority_chip has no runtime validator anywhere in the pipeline; the sister product has been observed emitting C-level, which is off-vocabulary, and it reached a pay band regardless. Both are open. Recorded 2 September 2026.

What is comparable today

FieldVocabularyReproducedComparable?
Seniority band (pay)closed, 5yesyes
Family · geographyclosedyesyes
Likely offer · ask · rangederived from those threeyes, to the euroyes
Pay confidence tierclosed, 3yesyes
Role archetypelisted, not lockedno — 3 of 5NO
Secondary archetypelisted, not lockedno — 1 of 6NO
Gap typenot closednoNO
Prose, ordering, phrasingfreeno, by designnot intended to be

So the honest claim is that two Role Intel briefs are comparable on money and on the classifications that set it, and not on the single word that names the role's shape. That word is the most quotable thing on the brief and it is the least reproducible.

What is confirmed, and what is inferred

Everything comes from the job description submitted. Nothing is looked up: the posting is not fetched, the employer is not researched, no listing is retrieved.

Confirmed from the document

  • Stated title, location, work model
  • Stated responsibilities and requirements
  • A published salary range, where the JD gives one
  • Stated team size, reporting line, budget
  • Named tools, stacks and certifications

Inferred by the system

  • Role archetype and secondary archetype
  • True seniority and scope versus the title
  • JD-inflation signals and severity
  • Occupational family and geography band
  • Pay figures, which follow from those two
  • Transferability and talent-pool depth

These are readings of a posting, not facts about the employer.

What the system does not do

  • It does not fetch the job posting. A brief that claims it retrieved the JD is blocked before delivery — that is a named gate rule.
  • It does not evaluate a person, and it never sees one. A Role Intel brief is a read of a document about a job.
  • It does not tell you whether to apply. It has no verdict field of that kind.
  • It does not price above a figure the employer published. The internal table may never outbid a published ceiling — also a gate rule.

The delivery gate

Deterministic code inside the renderer, not a model. Nineteen numbered rules, DG-01 to DG-19, gate version 1.8. Each finding carries a rule id, a severity, a message and a machine-readable evidence object. Any finding at BLOCK stops delivery; WARN findings are recorded and the brief ships.

Why 1.8, briefly: 1.7 could reject its own correct output. One rule demanded a verdict — under-priced, over-priced, or aligned — on every brief, including ones priced against a job description that published its own salary grid, where no verdict was ever computable. 1.8 tells the two cases apart before asking for a verdict. The incident that surfaced it is below, dated.

#Checks#Checks
DG-01Indexability / robots contractDG-11Fabricated provenance ("retrieved the JD")
DG-02Verbatim JD span leakageDG-12Renderer produced a deliverable brief
DG-03Third-party email / phone leakageDG-13Function-area count vs calibrator
DG-04Named third-party employers on a public pageDG-14No figure above a published ceiling
DG-05Salary node contract completenessDG-15Ask basis matches the ask figure
DG-06Scope bump must not leak into likely offerDG-16Floor-only and degenerate disclosed ranges
DG-07Estimated-family band above Estimated confidenceDG-17Unanchored geography vs a rendered figure
DG-08EU anchor cap against evaluated valueDG-18Editorial claims vs evidence
DG-09Offer plausibility, disclosed and modelled pathsDG-19Comp scenarios vs the offer figure
DG-10Banned strings in the delivered brief
The gate reports everything it runs — on briefs that ship. On blocked briefs it reported almost nothing, until today. A delivered brief carries delivery_gate_findings: the full serialised list of every finding, with ids, severities and evidence, plus a pass flag and block and warn counts. But a blocked brief throws, and the throw replaces the output — so the findings were lost precisely when they mattered. On 1 September a real brief was blocked and the stored reason was 77 characters of the tail of one sentence, with every rule id cut off; identifying the rule took a reconstruction from the payload. Two fixes have shipped: the message is now a single line, and the rule ids are appended at the end, where the truncation keeps them. Recorded 2 September 2026.

Determinism — what we measured

Same intent as the column beside this one, same answer in shape, but the split falls in a different place: Role Intel's money reproduces and Role Intel's headline word does not.

Identical input · pay path

Two runs of the same job description on the same day reproduced the pay figures exactly. A third run, captured separately as a fixture on live code, reproduced an earlier brief's figures to the euro. A fourth case was checked by recomputing the arithmetic by hand from the published band: likely offer at the lower third of a €89,700–€152,300 grid gives €110,566.7, rounded to €110,550, and the ask at the grid maximum gives €152,300. Both matched the delivered figures digit for digit.

The pay engine is deterministic code and behaves like it. Where pay has moved between runs, the cause has been a classification changing upstream, not the table.

Identical input · the archetype

Sixteen runs of byte-identical input returned Builder six times, Operator six times, Fixer four times — near-uniform across three of the seven values. Temperature was tested at three settings including 0; pinning it made the spread worse, not deterministic.

Separately, six real job descriptions were each submitted twice by accident, hours apart, byte-identical. The dominant archetype reproduced in 3 of 5 comparable pairs, the secondary in 1 of 6. In one case a single submission was delivered twice from literally the same bytes and came back Operator/Translator once and Builder/Connector the other.

The cause is known and it is not mysterious: the seven archetypes are handed to the model as a bare list with no definitions. Defining them is the next change, and it is the cheapest experiment on the backlog. Nobody has run it yet.

A question for whoever maintains the Profile Intel column. The changelog below records an earlier sixteen-input reproducibility study as unrecorded because its date and fields cannot be sourced. The sixteen-run study above is a Role Intel archetype study. These may be the same study, filed against the wrong product. We are flagging the possibility rather than merging them, because guessing would be exactly the failure the log exists to prevent.

What we claim, precisely

  • The generation call runs at temperature 0.3, set explicitly at the request, and the reviewer at 0.3 as well.
  • Seniority band, family, geography and every pay figure reproduced exactly across the identical-input cases above. Evidence from a handful of documents, not a guarantee.
  • The role archetype does not reproduce, at 3 of 5 and 1 of 6, on samples that are small but consistent, and on a mechanism we can name.
  • Whether the generation call ever ran without an explicit temperature is not established. The retained workflow history reaches back only to 1 September 2026 — four versions, all from that day — so the question cannot be answered from it. It is not answered by assumption either.

EU AI Act posture

Not analysed, and not inherited. The Profile Intel analysis turns on a brief being commissioned by the person it describes. A Role Intel brief describes a job, is commissioned by a candidate reading that job, and evaluates no natural person at any point — a structurally different question that has not been worked through. The mechanism set out under the common engine below is stated there because it is the same whoever reads it; where the line falls for Role Intel specifically is open.

Submitted material and retention

Not established. Whether Role Intel makes the same retention promise as Profile Intel, and whether a mechanism stands behind it, has not been audited. Stating it before checking is how the Profile Intel promise ran ahead of its mechanism, and the same mistake is not worth repeating on this side of the page.

Delivery identity

A Role Intel brief declares product = "Role Intel", and that label appears in the brief's provenance paragraph and in its structured page data. Whether the two products are visually distinguishable to a reader who lands on one has not been audited.

Open backlog

  • Role archetype does not reproduce and the seven values are undefined in the prompt. 2 Sep 2026.
  • Gap type has no closed vocabulary — 8 labels across 2 runs. 2 Sep 2026.
  • No runtime validator on the seniority chip, in either product. 2 Sep 2026.
  • Thirteen non-monotonic salary rows: Senior Manager prices above Director in two families, so a scope adjustment upward lowers the band while the brief says it was adjusted for scope. 2 Sep 2026.
  • ic_junior has no band label and cannot be scope-adjusted; a junior role is priced correctly and captioned with the senior band's name. 2 Sep 2026.
  • Three phantom rungs are referenced by consumers and producible by none. 2 Sep 2026.
  • Pay figures are not stored, so the database cannot answer what a given brief told a customer about money. 2 Sep 2026.
  • Romanian briefs leak English interface strings — 59 sites inventoried, 46 of them reaching the page. The locale itself resolves correctly; these are untranslated constants. 2 Sep 2026.
  • The scope-adjustment direction is emitted and never read by the renderer, so a brief can argue a title is inflated and then advise as though it were understated. 2 Sep 2026.
  • The engine version stamped into the brief index is four versions stale. 2 Sep 2026.
Profile Intel €19 Measured 1 Sep 2026

A structured capability read of one person, from their CV, commissioned by that person.

Closed vocabularies

A closed vocabulary means the system may only emit values from a fixed list. It is what lets two briefs be compared. Where a vocabulary is not closed, this page says so.

Profile types 7, closed

Builder · Operator · Specialist · Connector · Fixer · Translator · Anchor

The prompt states these must be exactly this set. A profile carries a primary type and may carry a secondary. The type is separate from the career archetype — a free-text positioning label capped at four words, which may never be one of the seven types.

Career patterns 7, closed

Linear · Progression · Steady · Lateral · Crossover · Multi-track · Multi-context

A profile carries a dominant pattern and, where the shape is dual, a foundation pattern. The dominant is always the more distinctive of the two, never the more common one.

Capability ladder 4 bands, closed

Every brief carries four to six capability clusters. Each gets a band, an evidence confidence and a recency class. The band is not a judgement of the person; it is a statement about how much of the CV supports the claim.

BandRequires
DistinctiveDocumented market outcome and rarity evidence and confidence High and recency Current or Recent
ExpertConsistent multi-role evidence at confidence High — or Med only where the cluster has a quantified outcome anchor
DemonstratedTwo or more roles, no outstanding outcomes; confidence Med or Thin
EmergingLimited, recent or unvalidated evidence; mentioned once or implied only

Evidence confidence

High3+ anchors, 2+ roles
Med2 anchors, or one role
Thin1 anchor, or implied

Recency

Currentwithin 2 years
Recent2–6 years
Historic7+ years

The recency gate caps the top of the ladder, and only the top. If the newest supporting evidence for a capability is seven or more years old and has not been refreshed by recent work, that capability caps at Expert rather than Distinctive. It is never dropped below Expert on recency alone. A rare credential stays a rare credential even when it is old — currency of a capability and rarity of a credential are different axes, and the system does not collapse them.

Legacy labels. Briefs generated before the ladder was renamed carry Solid, Developing and Basic. These map to Demonstrated, Emerging and Emerging. They are not emitted in new output.

One vocabulary is not closed, and you should know which. The management-tier scale used in the experience timeline — IC / Team Lead / Manager / Senior Manager / Director / VP / C-level — is specified as a list but not locked the way the vocabularies above are, and its seven values do not map one-to-one onto the six-rung scale the pay engine uses. A fix is specified and not yet shipped. Recorded 1 September 2026; see the changelog.

What is comparable today

Comparability rests on the closed value lists, not on the prose being identical. So it holds for exactly the fields whose vocabularies are closed and which have been shown to reproduce:

FieldVocabularyReproducedComparable?
Profile typeclosed, 7not measuredby vocabulary
Career patternclosed, 7not measuredby vocabulary
Capability bandclosed, 4not measuredby vocabulary
Confidence · recencyclosed, 3 eachnot measuredby vocabulary
Seniority band (pay)closed, 6yesyes
Family · geographyclosedyesyes
Likely-offer figurederived from those threeyesyes
Management tiernot closednoNO
Peak direct reportsan integer, not a vocabularynoNO
Prose, ordering, phrasingfreeno, by designnot intended to be

So the honest form of the claim is narrower than "briefs are comparable". Two briefs can be compared on profile type, career pattern, capability band, seniority, family, geography and the pay figures. They cannot currently be compared on management tier or peak direct reports. Those two are named in the changelog as open, and this table is what the claim is reduced to until they close.

What is confirmed, and what is inferred

Everything on a brief comes from the submitted CV. Nothing is looked up about the person, and no external record is consulted. But not everything on a brief has the same standing.

Confirmed from the document

  • Employers, role titles, stated dates
  • Durations computed from those dates
  • Stated skills, tools, named credentials
  • Stated figures, as the CV states them
  • Education and certifications as listed

Every evidence anchor must trace to specific content in the CV. An empty anchor list is a correct output; an invented one is a failure.

Inferred by the system

  • Profile type and career pattern
  • Capability clusters, bands, confidence, recency
  • Seniority band and occupational family
  • Salary bands, which follow from those two
  • Rarity anchors and the misread-risk audit
  • Trajectory, fit and direction

These are readings of a document, not verified facts about a person.

What the system does not do

  • It does not verify that a claim in the CV is true. It audits how well the CV supports the claim, and marks each one substantiated, partial, weak, unverifiable or unsupported.
  • It does not look the person up. No search, no public records, no social profiles.
  • It does not rank one person against another. A brief is commissioned by, and about, one person.
  • An impressive company-scale number the audit cannot substantiate is treated as context, never as a lift to a band, a confidence, a level or a trajectory.

The delivery gate

Before a brief is delivered, a separate reviewer model checks it against numbered rules and returns pass, warn or fail on each. The gate owns the final decision. There are 29 numbered phases — eleven K phases and sixteen L phases, plus a truncation check and a language check. Six L phases carry lettered sub-checks. There is no K8.

#Phase#Phase
K1Version compatibilityL1Required fields
K2Seniority calibrationL2Type and enum validation
K3Calibration gatesL3Inflation-signal validation
K4Path coherenceL4Friction-point type enum
K5Type-confidence validationL5Inflation coherence
K6Market-adjacency completenessL6Recency rule
K7Unique-signal nullabilityL7CV-completeness mirror
K9Delivery gate (final call)L8Vocabulary gate, inflation risk
K10Substantiation validationL9Language consistency
K11Contamination resolutionL10Distinctive-edge structure
K12Macro ranking-implicationL11Value-layer enum
YTruncation checkL12Capability schema
ZLanguage gateL13Searchable skill tags
L14Visible / internal field contract
L15Source-citation defaults
L16Premium-eligible flag
The gate reports a subset of what it runs. All 29 phases execute, but the machine-readable result carries fifteen entries: thirteen rules with a pass / warn / fail, plus two warn-only sub-checks (seniority anchor, tenure coherence). The remaining sixteen phases act on the output and raise warnings, but do not surface a per-phase verdict. We are stating this rather than publishing a rule count that sounds larger than what is reported. Recorded 1 September 2026.

A brief is delivered on pass_validated or compact_validated. On fail_validated it is held and not sent. A held brief is a working gate, not a lost order — it means the reviewer found something the read could not stand on.

Determinism — what we measured

The design intent is that the same CV produces the same brief. That is not what we measured, so it is not what we claim.

Identical input, two runs · 1 Sep 2026

FieldRun ARun B
Seniority bandDirectorDirectorstable
Occupational familymarketingmarketingstable
GeographyRO_bucharestRO_buchareststable
Likely-offer figure€3,250€3,250stable
Inflation-risk labelMediumMediumstable
Market-value figure€4,017€4,095moved
Peak direct reports14moved
Coordination-scope labeltwo different stringsmoved

The classifications that set the salary reproduced exactly. The management-scope reading did not.

Eighteen characters changed · 1 Sep 2026

The same document with an eighteen-character phrase deleted from one bullet. Nine of twelve timeline rows were identical. Three moved, and the peak-level reading moved with them:

RoleWithWithout
Head of Marketing & GrowthDirectorManager
Marketing & Operations LeadDirectorManager
FounderC-levelIC

Peak management level for the whole profile: Director in one run, Manager in the other — a two-band swing on the same document. The claims audit also differed: twelve claims listed in one run, thirteen in the other, and the one claim marked unsupported in the first was not audited at all in the second. That single difference propagated to the inflation-risk label and to whether a whole block appeared on the page.

The prediction under test — that the phrase was setting the seniority band — was wrong. The band did not move; the management tier did. We are recording the failed prediction because it is the result.

What we claim, precisely

  • Generation runs at temperature 0.3. It is a language model reading a document, and it does not repeat itself exactly.
  • On the evidence above, seniority band, occupational family, geography and the likely-offer figure reproduced exactly across identical input. That is a measurement on one document with two runs — it is evidence, not a guarantee.
  • Management tier, peak direct reports and the claims audit did not reproduce. Named, dated, and open.
  • Both studies were run after 28 August 2026, so they describe the engine as it is configured today. They are not a measurement of the briefs produced before that date, which ran at a materially higher temperature — see the changelog.
  • Comparability rests on the closed value lists, and the table above names which fields that covers today.

EU AI Act posture

A Profile Intel brief is a capability overview commissioned by the person it describes, and in that use it is not a high-risk AI system. Using one to evaluate, screen, rank or compare someone else against a role is a different purpose and may fall within high-risk classification. The mechanism, and the date it applies, is set out under the common engine below.

Open backlog

  • Management-tier vocabulary is not closed. Fix specified, not shipped. 1 Sep 2026.
  • Peak direct reports does not reproduce, and it is load-bearing on the pay figure via a peak-management premium. 1 Sep 2026.
  • The claims audit varies in selection, not rating — which moves the inflation label and whether a whole pay block renders. 1 Sep 2026.
  • An earlier sixteen-input reproducibility study is unrecorded because its date and fields cannot yet be sourced. 1 Sep 2026.
  • Suppressed-pay wording renders in English on Romanian briefs. 2 Sep 2026.

The common engine

Both products run on this. It is stated once here rather than twice above, because a claim written in two places is a claim that will eventually disagree with itself.

Architecture

Deterministic code around model calls: a generation step that reads the document and writes the structured output, and a separate reviewer that checks that output against numbered rules before anything is delivered. Everything between and around them — extraction, normalisation, the pay lookup, the gate arithmetic, the rendering — is ordinary code and behaves the same way every time.

The number of calls is not the same on both sides, and this page said it was until 2 September 2026. Profile Intel runs two. Role Intel runs three: a small calibrator that classifies the posting first, then generation, then the reviewer. Model and temperature differ too — see the table beside this one.

Model configuration

Profile Intel

Generationclaude-sonnet-4-6temp 0.3, max_tokens 20,000
Reviewerclaude-haiku-4-5temp 0

Role Intel

Calibratorclaude-haiku-4-5temp 0.3
Generationclaude-sonnet-4-6temp 0.3, max_tokens 32,000
Reviewerclaude-sonnet-4-6temp 0.3, max_tokens 16,001

Temperature 0 on the Profile Intel reviewer is not determinism either — it is a decoding setting, not a guarantee of an identical result. Role Intel has no call at 0: all three run at 0.3, and the reviewer is the larger model rather than the small one.

The temperature boundary, and how much of the delivered corpus sits on the wrong side of it. Before 28 August 2026 the Profile Intel generation call set no temperature at all and ran at the API default of 1.0 — the only model call in the product without an explicit setting, while the reviewer beside it already sat at 0. It writes the whole brief, structured fields included. Counted against the engine fingerprint stored with each brief: 109 briefs were generated before the fix and 9 after it, with 13 produced too early to carry a fingerprint at all. Roughly five in six Profile Intel briefs delivered to date were produced at temperature 1.0. Briefs from either side of 28 August 2026 are not directly comparable artefacts, and the engine fingerprint recorded with each brief is what distinguishes them. Whether the Role Intel generation call had the same gap is not yet established, and cannot be settled from the workflow history: the retained history for the Role Intel workflow reaches back only to 1 September 2026 — four versions, all from that day's work. Role Intel's generation temperature is explicit today, applied to the request at call time. How long it has been is an open question with no evidence behind it either way.

What "deterministic" does and does not mean here

The pipeline is deterministic code around a language model, and the code is deterministic — the same inputs produce the same pay lookup, the same gate arithmetic, the same rendered layout. The classification step is not. Any claim that "the same input produces the same output" has to say which of those two it is talking about, and ours did not. That is why it was withdrawn rather than restated.

Submitted material and retention

Where the EU AI Act line falls

A brief commissioned by the person it describes is not, in that use, a high-risk AI system. Using one to evaluate, screen, rank or compare someone else against a role is a different purpose, and may fall within high-risk classification under Annex III point 4(a) of Regulation (EU) 2024/1689 — AI systems intended for the recruitment or selection of natural persons, in particular to evaluate candidates — via Article 6(2). Article 25(1)(c) is the route by which provider obligations transfer: a party who modifies the intended purpose of a system so that it becomes high-risk is considered its provider and takes on the obligations in Article 16. Those obligations apply from 2 December 2027, deferred from 2 August 2026 by Regulation (EU) 2026/1744 — so this describes where the line will fall, not an obligation in force today. It is a statement of scope, not legal advice.

This states the mechanism, which is the same whoever reads it. How it applies to Role Intel specifically has not been analysed — Role Intel reads a role rather than a person, which is a different question and does not inherit the Profile Intel answer.

On the roadmap

Not defects — features that don't exist yet, named here on purpose so they don't quietly drop off the list. Dated when work starts. Moved to the changelog, and off this list, when they ship.

Role Intel — show the independent estimate next to a published band

Today, whenever a job description publishes its own salary band, Role Intel still computes its own independently modelled value for the same work — and then doesn't show it, deferring entirely to the employer's number. The plan is to show both: what they publish, what the model independently values the work at, and the gap between the two, stated plainly.

The gap runs both ways, and the copy has to say so. Sometimes the independent estimate is higher than the published band — under-priced scope. It can also be lower: early-stage, well-funded roles in particular sometimes publish a band ahead of what the seat's actual scope would independently price at. Neither direction is the one to expect by default; the point is showing the number, not editorialising it.

Role Intel — define the seven role archetypes in the prompt

The archetype is currently handed to the model as a bare list with no definitions — see the determinism section above, where it is the field that doesn't reproduce. Writing the definitions is the leading hypothesis for why, and the cheapest experiment on the backlog. Not yet run.

Profile Intel — close the seniority vocabulary

Seven values are listed for management tier; only six are locked in the pay engine, and nothing checks the list at runtime today (see the open-vocabulary flag above). The plan is one shared enum across both consumers, hard-failed in deterministic code rather than judged by a model.

What this page shows, and what stays proprietary

Enough to check a claim on a brief against a stated definition, and to see the shape and pace of the work. Not enough to rebuild the engine.

Shown here

  • The full classification vocabularies — closed and not-closed alike
  • What the delivery gate checks, and what it reports versus what it runs
  • Measured determinism figures, including the ones that came out badly
  • A dated changelog of every claim added, changed, or withdrawn
  • Model and temperature configuration for both products

Not shown here

  • The generation and reviewer prompts themselves
  • How a value gets assigned to a vocabulary term — the classification logic
  • The pay-derivation tables and their weighting
  • The Role × Profile matching logic, once it exists

The line is drawn where it is on purpose: what's shown is what a reader needs to check a claim on their own brief. What's kept is what would let someone reproduce the read without doing the work of building it.

Changelog

Claims added, changed and withdrawn. Newest first. Every entry carries a date and a reason. A withdrawn claim stays on this page — removing it would defeat the point of the page.

2 September 2026
AddedThe Role Intel column is filled inRole
Vocabularies, the nineteen gate rules, measured determinism and the open backlog, all read from the live workflow rather than from notes. Six of the eight outstanding items are now answered. Two are answered with "not established" and stay that way: retention, and whether the generation call ever ran without an explicit temperature — the retained workflow history only reaches back to 1 September 2026, so there is nothing to read. The EU AI Act question is recorded as open rather than inherited from the Profile Intel analysis, because a read of a role is not a read of a person.
2 September 2026
Withdrawn“Deterministic code around two model calls”Both
Stated as shared architecture; it is not shared. Profile Intel runs two model calls. Role Intel runs three — a calibrator, then generation, then the reviewer — and its three run at 0.3, 0.3 and 0.3, where Profile Intel's reviewer runs at 0. A sentence written once to avoid two copies disagreeing had instead made one description cover a system it did not describe.
2 September 2026
Known gapTwo vocabularies are open on the Role Intel sideRole
gap_type has no fixed list at all — two runs of one input produced eight distinct labels with a single overlap. And role_seniority_chip has no runtime validator anywhere in the pipeline, in either product; the Profile Intel side has been observed emitting C-level, which is off-vocabulary, and it reached a pay band. Named here because the Role Intel column now claims a closed vocabulary for other fields, and that claim is only worth something if the exceptions are named beside it.
2 September 2026
AddedMeasured Role Intel determinism, including the part that failsRole
Pay reproduces: identical inputs returned identical figures, one case checked to the euro against hand-recomputed grid arithmetic. The role archetype does not: sixteen runs of byte-identical input returned Builder six times, Operator six times and Fixer four times, and six accidentally duplicated real submissions reproduced the dominant archetype in 3 of 5 pairs and the secondary in 1 of 6. Pinning temperature to 0 made the spread worse. The cause is named — the seven archetypes are given to the model as a bare list with no definitions — and the fix has not been tried.
2 September 2026
ChangedA blocked Role Intel brief now records which rule blocked itRole
A delivered brief has always carried the full serialised gate findings. A blocked one threw, and the throw replaced the output — so the reason vanished exactly when it was needed. On 1 September a real brief was blocked and the stored reason was 77 characters of the tail of one sentence, with every rule id cut away; identifying the rule required reconstructing the arithmetic from the payload. The message is now a single line and the rule ids sit at the end, where the truncation keeps them.
2 September 2026
ChangedThis page now covers both productsBoth
Restructured so Role Intel and Profile Intel sit side by side above the shared engine. The Role Intel column is published empty and dated rather than filled with an unverified summary: it lists the eight things that must be measured before anything is claimed there. Setting the two columns beside each other makes the gap visible, which is the point.
2 September 2026
AddedThe retention promise has a mechanism — and one half of it does notBoth
A scheduled clear of stored CV text and the file reference was built for Profile Intel. It is stated here with the part it does not deliver: the uploaded file itself sits with the form provider, and removing our reference to it is not the same as deleting it there. That step is not yet in place. Saying so is better than letting the promise run ahead of the mechanism a second time.
1 September 2026
Withdrawn“Pay figures are calibrated for the market shown”Profile
Salary bands are held per occupational family and geography. A location that did not resolve to one of the mapped regions fell through to the Central and Eastern European band and printed a real euro figure from it, under a one-line caveat that the region was not calibrated. A caveat above a specific number is not the same as declining to give one — the number is what a reader takes away. From today those profiles show no pay figure at all, and say why. This is forward-only. Briefs delivered before 1 September 2026 for unmapped locations carry a figure derived from a European band that was not calibrated for their market, and it stands on the page as it was published. If you hold one of those briefs and want it re-issued or removed, ask.
1 September 2026
AddedHow much of the delivered corpus predates the temperature fixProfile
Counted against the engine fingerprint stored with each brief: 109 briefs were generated before the fix and 9 after it, with 13 produced too early to carry a fingerprint at all. Roughly five in six briefs delivered to date were produced at the API default temperature of 1.0. An independent count by record date agrees to within one row. This matters twice over. Both determinism studies on this page were run after the fix, so they measure the engine as it is configured now and say nothing about how the earlier briefs would reproduce. And it applies to any downstream use of the corpus, including as evaluation or training material.
1 September 2026
AddedThis pageProfile
Published the vocabularies, the numbered gate rules and the measured determinism status in one place, so a claim on a brief can be checked against a stated definition rather than taken on trust.
1 September 2026
ChangedAI Act citation correctedBoth
Earlier wording cited Article 25 alone. Article 25 is the mechanism that transfers provider obligations to whoever repurposes a system; it is not what makes the system high-risk. The classification route is Annex III point 4(a) — AI used to evaluate candidates — via Article 6(2). Both are now cited, along with the date those obligations actually apply, which the earlier wording omitted and which materially changes what a reader should infer.
1 September 2026
ChangedGate-rule count stated as reported, not as runProfile
Twenty-nine numbered phases execute; fifteen entries are reported in the machine-readable result. Publishing the larger number without that distinction would overstate what the gate demonstrably checks per brief.
1 September 2026
Withdrawn“All classification vocabularies are closed”Profile
The management-tier scale in the experience timeline is specified as a list but not locked, and its seven values do not map cleanly onto the six-rung scale the pay engine uses. Found while investigating a two-band swing in the peak-level reading between two runs of the same document. A fix is specified. The claim comes back when the fix ships and a regression test holds it — narrowed in the meantime to the comparability table above, which names the fields the claim still covers.
1 September 2026
Withdrawn“The same input produces the same output”Profile
Withdrawn on the strength of two controlled studies run today. The first put the same document through the pipeline twice: seniority, family, geography and the pay figures reproduced exactly, but peak direct reports moved from 1 to 4 and the coordination-scope label came back as a different string. The second changed eighteen characters and moved the peak management level two bands, from Director to Manager, along with the length of the claims audit. The claim was made because the pipeline is deterministic code around a language model, and the code is deterministic — but the classification step is not, and the claim did not distinguish between them. Replaced by the measured status above, which names which fields hold and which do not.
1 September 2026
AddedMeasured determinism figures, with the failed predictionProfile
Two runs of an identical document, and an eighteen-character A/B on the same document. The A/B was designed to show a phrase was driving the seniority band; it showed the opposite — the band held and the management tier moved instead. Recorded as it came out.
1 September 2026
Known gapAn earlier study is missing from this logProfile
A controlled reproducibility study of sixteen identical inputs was run before the two above and found non-deterministic classification. It is not recorded here because we cannot yet state its date or name the specific fields it moved, and this page does not carry claims it cannot source. It will be added when both are established. We are naming the gap rather than leaving it out, because a changelog that quietly omits an earlier finding is worth less than one that says where it is incomplete.
28 August 2026
ChangedGeneration temperature pinned to 0.3Profile
Today the generation call runs at temperature 0.3 and the reviewer at 0. Before 28 August 2026 the generation call set no temperature at all and ran at the API default of 1.0 — the only model call without an explicit setting, while the reviewer beside it already sat at 0. It writes the whole brief, structured fields included. Briefs produced before this date were generated at a materially higher temperature than briefs produced after it, and are correspondingly less likely to reproduce. This is the largest known determinism change in the system, and it means briefs from either side of 28 August 2026 are not directly comparable artefacts. The engine fingerprint recorded with each brief is what distinguishes them.