Skip to main content
Lotlsoft Labs

What bootstrapped founders actually said about AI, pricing and getting acquired

How this was made

Six steps. The first three are what anyone with a research agent does today. The last three are the ones almost nobody does, and they are why 816 claims are not on this site.

If you only read one part, read the refusals. Four of them side by side is a better argument than anything on this page.


1. The corpus

Four podcast back catalogues, pulled as RSS.

Show Episodes
The Bootstrapped Founder 442
Startups For the Rest of Us 329
MicroConf On Air 283
Indie Bites 131
1,185

Spanning 2020-02-05 to 2026-07-29. No model involved: fetch, parse, normalise.

2. The cohort

Who does the corpus keep talking about? 326 recurring subjects, extracted deterministically, using the podcast namespace's own <person> tags where the feed provided them and title and description patterns where it did not. Every subject records how its name was found, so the extraction can be audited rather than trusted.

Ranked by how often and how recently they appear and how much evidence the corpus already carried about them. The top 40 went to research. The other 286 did not, which is a budget decision and not a judgement about them.

Still no model. Steps 1 and 2 are a script.

3. The research

This is the step you already do.

Ten agents, four subjects each, with web search and page fetch. Each was asked to establish the same seven things per subject: what they run, whether it is still operating, how it is monetised, the most recent revenue figure on the record, how they acquire customers, what they have concretely changed because of AI, and what buyers valued if they sold.

The brief was not casual. Every claim had to carry the verbatim text that established it and the URL it came from. Every claim had to carry an asOf date, because a revenue figure from a 2024 interview is a 2024 fact and not a current one. notFound was explicitly a good answer. Padding was explicitly not.

The agents took it seriously. Several verified their quotes character by character against raw page text rather than against a summariser. Several caught and discarded third-party figures that contradicted a founder's own words. One found that a widely repeated "$32.5M exit" was a growth investment, not a sale.

They returned 366 claims.

4. The problem with what came back

The 366 looked good. They were sourced, quoted, dated, and produced carefully.

They were also not checkable, because a single claim was carrying several facts at once:

"Runs SavvyCal, a meeting-scheduling SaaS he founded in early 2020 and still runs with a team of about four, sold to individuals and teams who share scheduling links."

That is four claims wearing one coat. Ask whether the cited quote supports it and the answer is neither yes nor no. The founding year might be supported and the team size not; the whole thing fails on its weakest part and takes the good parts with it.

So the claims were split into the separate facts they actually assert, each keeping the evidence the original cited. 366 claims became 1,416 atomic ones, an average of just under four facts per claim.

This step is unglamorous and it is the one that makes the next step mean anything.

5. The check

Every one of the 1,416 was asked a single question:

Do this claim's own cited quotes actually establish it?

Not "is this plausible". Not "does the quote look relevant". An on-topic, correctly transcribed quote does not license a claim it never makes.

600 survived. 816 did not.

Not because the agents were careless. Because being careful at the research step does not get you there. The failure is quiet by construction: a real quote sitting under a claim that says slightly more than the quote does.

What that ratio is, and is not

600 of 1,416 is 42%, and we would rather you did not carry that number away as a fact about AI research, because three things move it and only one of them is about the research.

The denominator was our choice. Splitting 366 claims into 1,416 was a judgement call. Split more finely and each fragment asserts something more specific, which makes it easier to refuse. A different splitting prompt gives a different pass rate on identical research.

It moved four times over while we fixed bugs in our own code. The first run passed 10%. Then 25%. Then 42%. Same claims, same evidence, three numbers, and the only thing that changed was bugs in our code: a machine-readable predicate slug leaking into the judge's input, and a bare URL where the judge wanted a described source. We have no particular reason to believe it has stopped moving.

We spot-checked, we did not audit. We read perhaps twenty refusals closely and they were right. We did not read the other seven hundred and ninety-six, and we never examined the other direction at all, the claims that passed and should not have.

So the defensible sentence is narrow: one judge refused 816 of 1,416 atomic claims, on a split we chose, with checking code we were still fixing. That is not an error rate, and it is not a measurement of anything general.

What is solid is the artifact. The 816 are published with the evidence that was supposed to support them and the reason each was refused, and you can disagree with any of them.

The 816 refusals are published here, grouped by how they failed, because the failures are the part that shows the check does anything.

How they failed

Refusals Failure mode
294 The quote says something adjacent, but not this
215 Scope creep: an unearned "all", "only", "biggest", "entire"
115 A number the quote does not contain
114 A date or period the quote never gives
109 The quote never names the subject it is being attributed to

The single clearest example carries the whole idea in one word:

Claim Rootabl originally had a single price point of $99 per month or $888 per year Quote "So it's a 99 a month or 888 a year." Refused the quote confirms the price points, but gives no evidence this was the original price. It only states the price as it was.

The numbers are right there. "Originally" is the whole overreach, and almost nobody reads past it.

6. The ledger

What survived was written into a claim graph where every fact keeps its quote, its URL, the date it was true, and how confident the researcher was. Conflicting values for the same fact are detected, and a number going up over time is recorded as evolution rather than filed as a contradiction.

Sources are registered with the domain they came from, so five articles repeating one wire story can be told apart from five independent ones.

Two source numbers appear on this site and they count different things. 1,185 is the corpus: the episodes step 1 pulled, which is what the cohort was extracted from. 204 is the number of distinct URLs the surviving claims actually cite. The second is the one that matters for checking a claim, and it is much smaller because forty subjects' evidence concentrates on a few pages each.

Steps 4, 5 and 6 are an engine called reliquary. We built it for a different product, pointed it at this research, and it immediately refused more than half of it. That was not the result we wanted and it is the reason this site exists: the refusals turned out to be more interesting than the survivors.


What this check does not do

It verifies that a claim follows from the evidence cited for it. That is all.

  • A grounded claim can still be false. If a founder misremembered a number on a podcast, the quote supports the claim and the claim is wrong. Grounding is not fact-checking.
  • A refused claim is not a false one. "Refused" means unproven by the evidence attached to it. Many of the 816 are probably true and were simply not sourced well enough to publish.
  • The people quoted did not make these claims. The sentences were written by research agents reading public statements. A refusal is a mark against our sourcing, never against the person quoted.
  • The cohort is self-selected twice over: people who agreed to talk publicly about their numbers, then ranked by how often they appear. Survivorship bias is heavy and it is not corrected for.

Doing it yourself

The pipeline is open. Each phase reads and writes plain JSON, so you can stop after any of them and look at what you have.

trawl acquire  <sweep>   # feeds in, normalised documents out
trawl cohort   <sweep>   # who does this corpus keep talking about
trawl tasks    <sweep>   # research briefs for your own agents
trawl merge    <sweep>   # fold the agents' output back in
trawl atomise  <sweep>   # one claim, one fact
trawl ground   <sweep>   # the check
trawl export   <sweep>   # what survived, and what did not

The interesting output is your own refusal list, not your pass rate. Read twenty of them and see whether you would have shipped the claims.