How a thesis becomes a record
Four steps, in a fixed order. Each one writes something that the next step can read but cannot change.
-
Observe
Collect sources, timestamp them, and measure the story against the company's real exposure to it.
-
Validate
Try to break the thesis with data as it stood at the cutoff, simple baselines and stated invalidation tests.
-
Decide
A named person approves, conditions or rejects it, for one stated use and time window.
-
Record
Hash, sign and store the record. Add outcomes later as new entries.
1. Observe
Most research starts from a story. FactorHelm starts by asking how much of that story the company actually earns from.
Signals monitors news and filings about 1,022 public companies drawn from the S&P 500, S&P MidCap 400, Nasdaq-100 additions and AI-adjacent names. For every story it measures narrative velocity — how far mention volume is above the company's normal rate — and catalyst age, because markets often react to old news presented as new.
Each company has a revenue-overlap profile: how much of its actual revenue is exposed to each kind of AI narrative. 151 profiles have been reviewed by a person against sources; 107 are provisional and cannot trigger a strict call; companies without a profile default to an overlap of 0.50, which is never low enough to pass.
Insights does the same for creator claims on Instagram and YouTube: it extracts the specific claim, its ticker, direction and horizon, checks for independent corroboration, and grades the creator against what happened later.
What gets written
An evidence snapshot: every source with its publication time, retrieval time, and content hash; every claim marked as a fact, an inference, a hypothesis or a model output; contradicting evidence kept beside supporting evidence; and an explicit information cutoff. Anything published after the cutoff is rejected.
2. Validate
The validation path is built to say no. It runs separately from the research that produced the thesis and cannot change it.
It verifies the snapshot's hash, rebuilds the market data as it was at the cutoff, and runs a registered recipe: an event study, a transparent baseline a person can explain, the thesis's own invalidation tests, and sensitivity checks on parameters fixed in advance. Every test it runs is counted, so a thesis can't be tested twenty ways until one looks good.
A more complex model only advances if it beats the simple baseline out of sample, after costs and leakage checks. Failed tests stay in the record.
The validation path is owned by the same people who build the research. We describe it as separate, not independent. Independent reproduction is something we will ask an outside reviewer to do, not something we claim.
3. Decide
A decision is made by a person, for a stated use, and it expires.
The decision states what was evaluated, which checks passed and failed, what is still missing, the scope it applies to, who approved it and until when. The outcome can be approved for a stated use, conditional, insufficient evidence, rejected, or expired. Changing the thesis, the scope, the evidence or the approver creates a new decision; the old one stays.
No decision grants permission to trade. Execution sits outside FactorHelm.
What agents may and may not do
Two agent roles are used, and both are bounded. The Evidence Scout looks for sources the thesis is missing. The Skeptic tries to reject it: it looks for contradicting evidence, leakage and alternative explanations.
An agent is kept only if it finds more useful evidence per reviewer-minute than a single-model or rule-based baseline, and never invents a source.
4. Record
The record is hashed, chained to the previous day's records and signed. After that, it can be read but not edited.
Each strict call's canonical record is hashed with SHA-256. Every UTC day's hashes go into one manifest that also contains the previous manifest's hash, so the days form a chain. The manifest is signed with an AWS KMS key and written to S3 Object Lock in compliance mode, where retention cannot be shortened, even by the account's root user.
Anyone holding the public key can check the chain offline, without trusting our database. Outcomes — a graded result, an invalidation, a correction — are new entries that point back to the record they concern.
The products
Signals
Narrative velocity, catalyst age and revenue overlap for 1,022 companies, with a decision evidence card for each: the gate matrix, an evidence timeline, source corroboration and an invalidation rule. Strict calls are anchored and graded publicly.
Insights
Creator claims from Instagram and YouTube, turned into specific, dated predictions, checked against independent sources and graded later. Each claim gets a research state: investigate, follow, wait or ignore.
Portfolio
Applies a written policy — a diversified core, bonds, and a small capped sleeve for research ideas — to produce a next-buy plan, then records what was actually executed at the broker. It has no broker connection and no language model in its decision path.
What the record lets you answer
- What did we know, and when did we know it?
- Which person, model or rule made each assertion?
- What evidence contradicted the thesis?
- Which tests were run, including the ones that failed?
- Who approved it, for what, and until when?
- What happened afterwards, and did it match what was expected?