SPEAKMAN.AI · Release v1.4.0

Google starts retiring Gemini 2.5 this month. Your install moves to 3.x on its next launch.

By David Speakman ·

SPEAKMAN.AI is the free, local, open source engine behind the same MCP workflows the hosted platform runs. Version 1.4.0 drops every Gemini 2.x model and moves the Gemini and Vertex AI defaults to the 3.x line, and the Windows build migrates your saved settings on startup so you're off 2.5 well before Google's shutdown dates. It also fixes a quiet bug in update_session that let downstream agents ignore the very change they were re-run to absorb.

speakmanai.log · example, first launch after upgrading INFO Migrated retired model default_model: gemini-2.5-flash-lite -> gemini-3.5-flash-lite INFO Migrated retired model standard_model: gemini-2.5-flash -> gemini-3.8-flash INFO Migrated retired model advanced_model: gemini-2.5-pro -> gemini-3.1-pro-preview

Each tier keeps its role and gets the current model for it. The config file is rewritten once, and later launches leave it alone.

3
model tiers moved from Gemini 2.5 to the 3.x line, for both Gemini and Vertex AI.
0
setup-page visits needed to keep an existing Windows install working.
8.4 → 8.4
compliance score before and after an architecture revision, pre-fix. The agent never noticed the change.
1
model now offered in every tier: gemini-3.8-flash.
Why This Release

A model retirement shouldn't be the thing that breaks your pipeline.

Every workflow agent asks for a tier (fast, standard or advanced) rather than a model name, and the engine resolves that tier to whatever your provider settings say. That design already meant no workflow file had to change for this release. What could still break was the settings themselves. On Vertex AI, Google retires the 2.5 models from public availability on October 20, 2026. A project that hasn't called a given model in the 90 days before that is blocked from it right away, and everyone else keeps access only until the shutdown dates in 2027. After that, a saved config pointing at gemini-2.5-flash fails on every call.

So the exe checks on startup. Any Gemini 2.x model in your saved settings is swapped for the new default of the same tier before the server comes up. Docker and run-from-source setups read models from .env, which the engine doesn't rewrite on your behalf. If that's you, copy the model lines from the updated .env.template.

TierWasNow
Fastgemini-2.5-flash-litegemini-3.5-flash-lite
Standardgemini-2.5-flashgemini-3.8-flash
Advancedgemini-2.5-progemini-3.1-pro-preview

gemini-3.8-flash is selectable in the Fast and Advanced dropdowns too, which matches Google's own guidance: its model versions page lists gemini-3.8-flash as a recommended replacement for all three 2.5 models, 2.5 Pro included. If 3.1 Pro is slower or pricier than your workload justifies, switching the Advanced tier is one click on the setup page.

The Bug Worth Explaining

The compliance agent had the new architecture in front of it and kept its old answer anyway.

update_session reopens a finished session, revises one agent, and re-runs everything downstream of it. On a real Solution Architecture run, the Tech Architect picked up signed QR payloads and a CCPA deletion module. The Compliance Officer re-ran and returned its previous report byte for byte, still flagging the deletion workflow as missing.

The event log showed the agent had received the revised architecture. The note sent with it was the problem. It named whichever step happened to run just before, never said what had changed, and closed with "otherwise leave your output unchanged." Given that, keeping the old report was a reasonable reading.

Before

"Upstream agent 'MCP_TECHNICAL_VISUALIZATION_SPECIALIST_V1' was just revised… otherwise leave your output unchanged." Compliance doesn't depend on that agent, and its output hadn't changed.

After

Downstream agents see which agent the update started from and the actual change request, along with an instruction to bring their output in line with the revised content.

What Shipped

The full list, not just the headline.

Models

Gemini 2.x removed from the setup page, the tier mappings, the templates and the docs. New 3.x defaults for Gemini and Vertex AI, with automatic config migration in the exe.

Revisions

update_session cascades now carry the origin and the change request downstream. Verified with a live re-run where the downstream output now changes.

Docker

render_document failed in the Docker image over a Python 3.12-only f-string construct. The exe's newer Python had hidden it.

Delegation

Submitting a delegate step's result could leave the session stuck showing IN_PROGRESS while it waited on the next step. The pause is now claimed atomically before the workflow resumes, and each response carries the pause_id it was written for, so it can only land on that pause.

The Docker and delegation fixes, along with a smaller review-gate hardening, were caught first in the hosted platform and ported back here.

Where This Sits

One engine. Free or hosted, it runs the same pipeline.

The hosted SPEAKMAN.AI platform runs a full idea-to-production SDLC pipeline. This repository is the free, local, MIT licensed engine that executes the same MCP workflows, with full functionality for an individual user and no account required.

Idea Intake

/generate-project-brief

A raw idea, coached into a structured brief. No SPEAKMAN.AI session, no credits.

Architecture

/generate-sad This release

Business description in, Solution Architecture Document out. Revising one section now carries through to the sections that depend on it.

Requirements

/generate-requirements

Domain model, naming dictionary, use cases, API and DB schema.

UX Design

/generate-ux-design

Screen inventory, flow diagrams, a proposed direction, and HTML mockups for the highest-priority screens.

Code Generation

/generate-speakmanai-code

A working, milestone gated, git committed codebase.

Infrastructure

/generate-infrastructure

Terraform for GCP, AWS, and Azure, plus pre-pentest hardening.

Pentest & UAT

/pentest

A gated security engagement, run against the live staging build.

Runs On Your Machine, From Whatever Talks MCP

The free tier is not a trial. It is the whole engine.

ModeCommandDatabase
Windows exeRun SpeakmanAI.exe~/.speakmanai/speakmanai.db
Docker + SQLitedocker-compose up --buildspeakmanai_data volume
Docker + MongoDBdocker-compose --profile mongo upMongoDB, port 27017
Dev, no Dockeruvicorn server:app --port 8000~/.speakmanai/speakmanai.db

There's no account and no cloud dependency, and you pay nothing per call beyond your own provider key. The Windows build is code-signed, so it installs without a SmartScreen warning.

Individual · Free · MIT

Run the engine locally.

Download the signed exe, or clone and run with Docker. You get full functionality on your own API keys, with nothing sent anywhere else.

Teams · Hosted

Need multi-tenant support and audit trails? Metered billing too?

The hosted SPEAKMAN.AI platform runs the same engine with the infrastructure teams need on top of it.

Written by David Speakman. Speakman Consulting designs and builds this kind of system for growing organizations: agent workflows with the governance that keeps humans in the loop.