I built a small LLM benchmark around false closure — the failures are rarer than I expected, but more interesting | TickerVault