We’re having great results from the Chatbase chat engine on fintechbenchmark.com. But we have to look ahead, and smaller models have caught our attention.
For years the industry assumed bigger was always better, with companies building ever-larger data centres to house models with trillions of parameters. But a shift is under way towards smaller, more efficient models, and for good reason.
Our website fintechbenchmark.com is a large directory of fintech products and services. The information on the site was seeded from the internet, and in this dynamic industry things change. To keep it current, we run an agent whose job is to keep the site up to date. It runs overnight, processing a batch of companies and their products, using AI to check that the information is accurate — and updating the site when things have changed.
I’ve written before about my reservations with Git. This is the story of how those reservations played out in practice — and what I should have done differently.
This week at the bridge club we had a minor crisis. We have a system for collecting scores each round and feeding them into a computer, which then produces the rankings. This week it lost both — the scores and, more importantly, the rankings. I say “more importantly” because we came top. Others may be less exercised about this.
As the club techie, it fell to me to sort it out. The results collection system works like this: small gadgets sit on each table for entering scores, these communicate wirelessly with a dedicated server, which is linked to a PC. A program talks to the server and creates a results file on the PC. So that file was the obvious first port of call.
Since my last post I’ve been using Claude Code more, and I’m pleased to report it works well and is genuinely powerful. As a confirmed control freak, I was also pleased to find that I stayed in control throughout.
After spending a year building fintechbenchmark.com the traditional way—writing code, testing it, iterating—I’d grown comfortable with GitHub Copilot as my AI assistant. It delivered a significant productivity boost, though I’d learned when to trust it and when to ignore its suggestions and push forward on my own.
Then, weeks before our soft launch, our lead contractor dropped a demo of Google Antigravity (Project IDX) that made our entire workflow look like stone tools and campfire stories.
The assumption that bigger AI models always deliver better results is being quietly dismantled — and the economics behind this shift are compelling. Now we have Small Language Models (SLM) to distinguish them from Large language Models (LLM). How do we measure the size of a model?
When I was seeding the database for fintechbenchmark.com with Claude-generated or Openai-generated content, I needed structured data that matched my database schema. My format of choice was JSON, and my approach was straightforward:
Request JSON in the prompt
Provide an example of the desired structure
Parse Claude’s response—a text string containing JSON wrapped in markdown code blocks with triple backticks
MCP (Model Context Protocol) solves a fundamental problem: how do you give an AI model access to external data like databases, files, or APIs without embedding everything in each request? Introduced by Anthropic in November 2024, MCP is now mature enough (15+ months old) for serious production use. Think of it like a USB-C port for AI—a standardized way to connect AI models to external systems.