By David Speakman ·
SPEAKMAN.AI is the free, local, open source engine behind the same MCP workflows the hosted
platform runs. Version 1.4.0 drops every Gemini 2.x model and moves the Gemini and Vertex AI
defaults to the 3.x line, and the Windows build migrates your saved settings on startup so
you're off 2.5 well before Google's shutdown dates. It also fixes a quiet bug in
update_session that let downstream agents ignore the very change they were
re-run to absorb.
Each tier keeps its role and gets the current model for it. The config file is rewritten once, and later launches leave it alone.
gemini-3.8-flash.
Every workflow agent asks for a tier (fast, standard or
advanced) rather than a model name, and the engine resolves that tier to whatever
your provider settings say. That design already meant no workflow file had to change for this
release. What could still break was the settings themselves. On Vertex AI, Google retires the 2.5 models from public
availability on October 20, 2026. A project that hasn't called a given model in the 90 days
before that is blocked from it right away, and everyone else keeps access only until the
shutdown dates in 2027. After that, a saved config pointing at gemini-2.5-flash
fails on every call.
So the exe checks on startup. Any Gemini 2.x model in your saved settings is swapped for the
new default of the same tier before the server comes up. Docker and run-from-source setups
read models from .env, which the engine doesn't rewrite on your behalf. If that's
you, copy the model lines from the updated .env.template.
| Tier | Was | Now |
|---|---|---|
| Fast | gemini-2.5-flash-lite | gemini-3.5-flash-lite |
| Standard | gemini-2.5-flash | gemini-3.8-flash |
| Advanced | gemini-2.5-pro | gemini-3.1-pro-preview |
gemini-3.8-flash is selectable in the Fast and Advanced dropdowns too, which
matches Google's own guidance: its model versions page
lists gemini-3.8-flash as a recommended replacement for all three 2.5 models,
2.5 Pro included. If 3.1 Pro is slower or pricier than your workload justifies, switching the
Advanced tier is one click on the setup page.
update_session reopens a finished session, revises one agent, and re-runs
everything downstream of it. On a real Solution Architecture run, the Tech Architect picked up
signed QR payloads and a CCPA deletion module. The Compliance Officer re-ran and returned its
previous report byte for byte, still flagging the deletion workflow as missing.
The event log showed the agent had received the revised architecture. The note sent with it was the problem. It named whichever step happened to run just before, never said what had changed, and closed with "otherwise leave your output unchanged." Given that, keeping the old report was a reasonable reading.
"Upstream agent 'MCP_TECHNICAL_VISUALIZATION_SPECIALIST_V1' was just revised… otherwise leave your output unchanged." Compliance doesn't depend on that agent, and its output hadn't changed.
Downstream agents see which agent the update started from and the actual change request, along with an instruction to bring their output in line with the revised content.
Gemini 2.x removed from the setup page, the tier mappings, the templates and the docs. New 3.x defaults for Gemini and Vertex AI, with automatic config migration in the exe.
update_session cascades now carry the origin and the change request downstream. Verified with a live re-run where the downstream output now changes.
render_document failed in the Docker image over a Python 3.12-only f-string construct. The exe's newer Python had hidden it.
Submitting a delegate step's result could leave the session stuck showing IN_PROGRESS while it waited on the next step. The pause is now claimed atomically before the workflow resumes, and each response carries the pause_id it was written for, so it can only land on that pause.
The Docker and delegation fixes, along with a smaller review-gate hardening, were caught first in the hosted platform and ported back here.
The hosted SPEAKMAN.AI platform runs a full idea-to-production SDLC pipeline. This repository is the free, local, MIT licensed engine that executes the same MCP workflows, with full functionality for an individual user and no account required.
A raw idea, coached into a structured brief. No SPEAKMAN.AI session, no credits.
Business description in, Solution Architecture Document out. Revising one section now carries through to the sections that depend on it.
Domain model, naming dictionary, use cases, API and DB schema.
Screen inventory, flow diagrams, a proposed direction, and HTML mockups for the highest-priority screens.
A working, milestone gated, git committed codebase.
Terraform for GCP, AWS, and Azure, plus pre-pentest hardening.
A gated security engagement, run against the live staging build.
| Mode | Command | Database |
|---|---|---|
| Windows exe | Run SpeakmanAI.exe | ~/.speakmanai/speakmanai.db |
| Docker + SQLite | docker-compose up --build | speakmanai_data volume |
| Docker + MongoDB | docker-compose --profile mongo up | MongoDB, port 27017 |
| Dev, no Docker | uvicorn server:app --port 8000 | ~/.speakmanai/speakmanai.db |
There's no account and no cloud dependency, and you pay nothing per call beyond your own provider key. The Windows build is code-signed, so it installs without a SmartScreen warning.
Download the signed exe, or clone and run with Docker. You get full functionality on your own API keys, with nothing sent anywhere else.
The hosted SPEAKMAN.AI platform runs the same engine with the infrastructure teams need on top of it.
Written by David Speakman. Speakman Consulting designs and builds this kind of system for growing organizations: agent workflows with the governance that keeps humans in the loop.