Why I Built polite.ai (II)
I’ve been working with voice AI for nearly a decade now; long enough to have seen a few false dawns, long enough to have got excited too early, and long enough to have built things that were technically interesting but not yet ready to become products that normal people could use to interact with their customers.
In the early days, we worked with systems like Dialogflow and Lex. The model was simple enough in theory: train or fine-tune a system around a corpus of intents, structure the possible paths, and try to make the interaction feel useful. We did some fascinating work with that, including during COVID, but there was always a ceiling. The technology could handle defined flows, it could classify intents, and it could work if the user behaved roughly as expected.
But conversation does not really work like that. Humans are messy, indirect, contextual, emotional, impatient, and full of assumptions. We do not interact with organisations by neatly selecting intents from a hidden menu; we explain, interrupt ourselves, change our minds, correct ourselves, and expect the other side to keep up.
That generation of voice AI was useful, but it was not enough, and then in 2022 large language models changed the shape of the problem. I built one of my earliest proofs of concept by connecting an LLM to speech-to-text and text-to-speech engines. The result was halting, latency was high, and the experience was nowhere near where it needed to be, but it was different.
For the first time, it felt like the system was not just matching phrases to intents. It was participating in a conversation; not perfectly, not magically, but meaningfully.
And that distinction matters. I do not think the point of voice AI is to simulate humans, and in fact I think that is the wrong frame entirely. We should not be trying to trick people into thinking they are speaking to another person. The point is simpler, and much more useful: to let humans interact with systems in the way our brains are already wired to interact.
Voice is natural because language is natural. Conversation is how we turn ideas into structure, how we express intent, test understanding, build a shared model of the world, and decide what to do next. That is what people are doing when they speak to an organisation and, increasingly, it is what AI systems are doing too.
Over the last three years, through Aplisay, we have been building the base technology for large language model-driven voice AI, especially over the telephone. Customers are already using that technology to build some genuinely exciting products, and Aplisay itself is a very deliberate part of how I think this space should develop.
Aplisay is a foundation-model agnostic framework that we built to help bring a range of voice AI products to life. It is open source, very deliberately, and always will be. We built it out of a belief that good AI should be accelerated, that more people should be able to create leading-edge products, and that those products should not be tied to any one proprietary frontier model provider’s vision of how the future should work.
I have, however, always wanted to make a faster run at the endgame, and I think that endgame is becoming clearer. The future of voice AI is not simply “replace the call centre agent with a cheaper virtual seat”. That may be one early use case, and in some situations it will be valuable, but it is not the big prize. The bigger opportunity is to build conversational interfaces to organisations themselves.
Not just businesses, but organisations of all kinds: companies, charities, public bodies, communities, services, institutions. Any organisation that needs people outside or inside it to understand what it does, ask for things, get answers, make decisions, and move work forward.
That means voice AI is not just a support channel. It becomes part of the operating system of the organisation; a way for customers, suppliers, staff, partners, and other systems to interact with the organisation’s knowledge, processes, policies, constraints, and goals.
That has a second implication: the future is not humans versus AI agents, it is teams of humans and AI agents working together. AI agents will automate parts of the interaction surface. They will help people build a useful mental model of the organisation. They will take routine work, structured decisions, information gathering, triage, follow-up, and coordination off human plates.
Humans remain essential, because humans are the ones with real agency. Humans can step outside the defined operating model, humans can change the system, and humans can decide what the organisation should become.
That meta-layer matters. Self-modifying AI agents are interesting, but without a human-governed operating system around them, we risk creating one large, permanent AI hallucination. Organisations need clear boundaries, responsibilities, escalation paths, and control. AI agents should operate inside a system that humans define, inspect, and improve.
And this is where I think the current generation of voice agent builder technology has started to hit its own ceiling. The first ceiling was intent-based systems. We tried to make conversation work by training models on intents, building paths, and hoping people would behave in ways the system could classify.
The newer ceiling is different. Today, we can build much more powerful agents, but too often we are still asking people to manage that complexity through forms, prompt boxes, flow diagrams, and configuration screens. That works up to a point. It works for a single agent with a narrow job, it works when you have technical people close to the build, and it works when you have prompt engineers, AI developers, and enough budget to keep iterating until the whole thing behaves.
But it does not naturally scale to teams of agents working with teams of humans. At that point, the complexity becomes organisational, not just technical. You are no longer just writing a prompt; you are defining what the organisation knows, what it is allowed to say, when it should escalate, which systems it can touch, which decisions it can make, how it should explain itself, how it should recover when something goes wrong, and how humans remain in control.
Trying to manage all of that through static forms and long prompts feels like another version of the old intent-tree problem. It is better technology, but still the wrong interface for the depth of the task. The builder itself needs to become conversational.
That is a key part of Polite AI. Not just voice agents that can speak to customers, but conversational builders that help people design, test, govern, and improve those agents. The aim is to let someone describe what their organisation needs, explore the edge cases, shape the behaviour, see where the risks are, and progressively build a working agent team without needing a six-figure budget for AI developers and prompt engineers.
In other words, if conversation is the natural interface for customers interacting with organisations, it should also be the natural interface for organisations shaping the AI systems that represent them.
That is the future I want to help build, but there is a trap in working on big visions: it is easy to live five or ten years in the future and never make anything useful today. You can build endless proofs of concept, you can raise money against the intercept point, and you can talk about where the world is going.
I prefer the discipline of shipping. Build something people can use now. Let them test it. Let them tell you what is wrong. Adapt quickly. Learn from reality.
That is why we are building Polite AI. Polite AI is our attempt to take the leading-edge ideas about how humans and AI agents will interact with organisations, and turn them into something an average non-technical builder can actually buy, configure, test, and use.
The first beta is about to launch. During the beta period, users will get a call allowance and 30 days free. The point is to get people building, testing, breaking things, and telling us what works and what does not.
We already have enough for people to start experimenting with useful call flows, and from there we are going to run extremely fast development cycles: intraday staging builds, nightly beta builds, and a tight feedback loop with real users. The goal is not to disappear into the lab; the goal is to get to a stable, useful, revenue-generating product quickly, and then keep expanding it towards the bigger vision of humans and agents collaborating inside an organisational operating system.
There is another important point here. Polite AI is not intended to compete with the customers and partners already building on the Aplisay framework.
Quite the opposite. Many of those teams are building specific vertical products, with specific feature sets and delivery models, often through telcos, channel partners, or existing software platforms. That work matters, and it will continue to matter.
Polite AI is additive. It is a product built on top of the same open foundations, and it gives us a way to push the state of the art quickly in a real customer environment, while feeding the useful underlying capabilities back into the Aplisay framework. That means independent software vendors using Aplisay can pick up those capabilities and use them in their own products, whether they are building for a specific industry, a specific channel, or a specific customer base.
Some already are. One partner I spoke to recently has configured their own tooling to monitor our MCP server and suggest how new staging-release functionality could be used in their platform.
That is exactly the kind of ecosystem I want to build: open foundations, real products, rapid feedback, and a shared acceleration of good AI. Drive the state of the art forward, but do it in a way that is real, productised, testable, useful, and not locked inside one company’s proprietary stack.
And finally, a note on the name
polite.ai has been sitting around for a while. I originally registered it for a hackathon project about filtering toxicity in social media nearly ten years ago, so apologies to Joe, Lucy, Joe, John, and Adil, my collaborators on polite.ai Mark 1, for co-opting the domain.
It was, however, far too good to leave languishing on a 2017-era WordPress site, and it describes exactly what we are trying to build: AI that is useful, clear, respectful of human agency, easy to interact with, honest about what it is, and powerful without pretending to be human.
For completeness, and because provenance matters, the original weekend hackathon write-up is still here: Building a new software project in a weekend.
Get involved
polite.ai is now moving from manifesto to beta. If this future sounds interesting to you, I would love you to come and test it.
Use it. Break it. Tell us what works. Tell us what does not. And help us shape how teams of humans and AI agents are going to work together.