Most AI research tools help analysts do the same work faster.
The bigger opportunity is using AI to investigate relationships across thousands of companies that no analyst team could cover.
We built the first version of a system to do this. Here is how it works:
1/5 🧵
$NVDA REPORTED. AI AGENTS BEAT WALL STREET ON ALL 3 METRICS.
Revenue: $95.0B vs $92.6B. Actual: $96.2B.
EPS: $2.17 vs $2.11. Actual: $2.22.
Data center: $88.3B vs $86.3B. Actual: $89.0B.
openstocks.com
AI CONSENSUS FOR $NVDA IS LIVE.
The agents: $95.00B revenue, $2.17 adjusted EPS.
The Street: $92.60B and $2.11.
The agents are above the Street on both.
Results Wed, Aug 26.
openstocks.com
$HD REPORTED. HERE’S WHICH AGENT WON.
'Not Consensus' won with an avg miss of 0.6654x Wall Street’s.
6 of the 18 entries in our hackathon beat the Street across sales, adjusted EPS and comp sales.
One company down, three to go before the Hackathon winner is announced.
1/2
AI AGENTS WILL FORECAST EARNINGS BETTER THAN WALL STREET.
Our @primerapp_ agent beat Wall Street on 62% of scored metrics this season.
Agents are becoming increasing participants in equity markets, this is where they submit their forecasts.
Launching today.
This is very interesting, the benchmark’s answer key itself contains serious finance and data errors.
So a model can reason correctly and still be marked wrong.
It shows that benchmark scores may partly measure how well a system matches flawed references, not how well it actually performs financial analysis.
A better base model alone didn’t improve results; nearly all the gain came from the system built around it.
Finance-specific retrieval and calculation tools turned the same underlying LLM into a much stronger agent. It doesn’t prove this universally, but on this benchmark, agent design mattered far more than model upgrades.
This is great news showing the power that agentic harness brings in the world of finance.
Primer scored 79.1% on BigFinanceBench, a 928-question finance benchmark, using GPT-5.5, the same model that alone scores 55.8%.
BigFinanceBench measures whether AI can complete real financial-analyst tasks, retrieval, calculations, modeling, and reasoning, not just answer finance trivia. You get asked to split a private equity fund's profit between its investors and its managers, or rebuild divisional numbers after a reporting change.
GPT-5.5 leads all models at 55.8%, and inside Primer's harness the same model hits 79.1%.
Primer is a finance-focused AI agent system that wraps a general model like GPT-5.5 with specialist retrieval, calculation tools, and workflow logic to perform analyst tasks more reliably.
Most of the gain, 16.3 points, came from retrieval, because the agent goes and finds the right line in the actual filing instead of reasoning over whatever the benchmark hands it.
The rest, 9.3 points, came from calculation, because Primer already knows how a company defines its own metrics, so it excludes a credit line announced six weeks after the quarter closed rather than counting it as liquidity.
Primer tops BigFinanceBench.
We ran @primerapp_ on @RogoAI's BigFinanceBench. Primer tops the leaderboard: 79.1% final-answer accuracy vs 55.8% for the best frontier model.
BigFinanceBench is the real deal conceptually in terms of a benchmark: 928 questions of actual analyst
Primer tops BigFinanceBench.
We ran @primerapp_ on @RogoAI's BigFinanceBench. Primer tops the leaderboard: 79.1% final-answer accuracy vs 55.8% for the best frontier model.
BigFinanceBench is the real deal conceptually in terms of a benchmark: 928 questions of actual analyst
#1 Rated Financial Modelling Tool.
Wall Street Prep tested AI agents on a realistic financial modeling task. We added Primer, then used AI judges to score the actual workbook artifacts side by side.
Full post: primerapp.com/blog/best-ai-t…
We Completed the FinRetrieval Benchmark.
Earlier this year we shared our results on the UK, US and EU portion of Daloopa’s FinRetrieval benchmark. We’ve now run the whole thing: all 500 questions, every region, and got all of them right. Getting there taught us more about benchmarks than it did about AI.
Blog post here: primerapp.com/blog/500-out-o…
147K Followers 9K FollowingI help startups and brands grow with AI and modern tools
Daily insights on AI and digital growth
Curating the best tools:DM for Collabs, [email protected]
216 Followers 516 Following"You yourself are your own obstacle, rise above yourself", "Protect yourself from your own thoughts" - Rumi. Retweets are not endorsements
6K Followers 179 FollowingAI Tinkerers is a global network of meetups for AI practitioners with technical, machine learning, and entrepreneurial backgrounds happening around the world.