All posts

Why hypotheses and findings exist

An agent can read your ad results, but what does it know about why you ran those ads? This is the problem behind hypotheses and findings.

One of the complications of having agents work on ad creative is that access to the results doesn't give them all the context they need. An agent can see that one ad got more clicks than another, and it can offer a plausible explanation. But does it know what you were trying to find out? Does it know that you tried something similar last month, or that the ad with the cheaper clicks brought in people who weren't interested in buying?

Some of this context lives in the metrics, but quite a bit lives in someone's head, a Slack thread, or a document. The problem gets more obvious when an agent forms an opinion during a session and the next session has no record of it. You can keep generating and testing ads without accumulating much understanding of why any of them worked. This is the problem we built hypotheses and findings around in Growth Chicken.

A hypothesis is a belief about your creative, e.g. "founder-led voice beats brand voice for cold traffic". It belongs to a project and can be tied to an audience. The intention is that you keep one hypothesis for that belief as you test it over time, rather than create a new one every time you run a batch of ads. A finding is what you conclude from one of those rounds, recorded under the hypothesis with a positive, negative, or neutral signal.

Findings are append-only. If you change your mind, you write a new finding that supersedes the earlier one. This adds a little friction, but being able to see how a belief developed is useful context in itself. "We thought this worked, then tried it with a different audience and it didn't" gives the next person (or agent) more to work with than an edited conclusion that only says it doesn't work.

There are a few ways to add context around these. An experiment can group the ads behind a finding if you ran a trial, though a finding doesn't require one. Hypotheses can also link to each other when one revisits, contradicts, refines, or supersedes another. These relationships matter because beliefs about creative rarely develop in isolation; a result might change what you think about an audience as much as what you think about a particular ad.

All of this is readable and writable through the API. There's a brief endpoint that gives an agent the current state of the project's hypotheses and findings before it generates anything. The loop we use is fairly simple: read the brief, make ads, run them, and write back a finding. The next session can start with that context instead of having to reconstruct it.

One decision here that I think matters: we don't compute findings from ad metrics. A person or an agent has to write the conclusion, and it can disagree with what the numbers seem to suggest. "CPC was better, but the leads were junk" is a useful finding. Automatically calling that ad a success because its clicks were cheaper would lose the part of the result that should inform the next decision.

Of course, writing a belief down doesn't make it correct, and giving an agent more context doesn't mean it will use that context well. What happens when an early finding is wrong and every later session treats it as settled? How much evidence should it take to revisit a belief? We have a place to keep that history now; figuring out how to reason well from it is still an open question.

Keep reading

Art Styles

We added ten art styles to the ad form. A few thoughts on what happens when you ask an image model to change how something looks.