September 2026
We’ve completed the first version of JudgementBench, a benchmark where models judge the quality of AI-generated macrostrategy-relevant ideas. At the time of writing, JudgementBench seems far from being saturated by frontier LLMs, and we’ve developed a scaffold that significantly outperforms the best unscaffolded models on it. We’ll continue to refine and extend the dataset.