We(a team of ex-EY employees) built FAB - Finance Agents Benchmark, testing AI Agents on the work behind financial due diligence. Agents can find the relevant facts, but still struggle to carry them through to a complete, reliable analysis.
We’ll keep expanding FAB to more companies and testing more models. Building and running this benchmark isn’t cheap, so we’re scaling it in stages.
The benchmark is public.
GitHub: Hugging Face: https://github.com/SecondState-ai/finance-agents-benchmark data room and tasks: https://huggingface.co/datasets/secondstate/finance-agents-b...
Found myself staring at nothing while working and prompting so I thought I could make my agents "look alive" at least
My partner and I have different organisation styles. I like surprises and to play things by ear where as they like to carefully plan thing out. So I built this to help bridge the gap a little.
Possibly this is of interest to very few other folks but since it doesn't really cost me anything to run I figure I may as well put it out there in case it helps anyone else :)