09 / Software & automation
AI features you can measure, not just demo
We build LLM-backed features together with the things that decide their fate in production: an evaluation suite, tracing on every call and cost under control.
A demo built on a language model impresses after two days of work. The trouble starts later: nobody can say how often the assistant answers well, what it does with a question outside its scope, or what a month of it costs. A feature you cannot evaluate isn't a product. It's a liability that someone eventually discovers in front of a customer.
Wiring up a model API is something almost anyone can do now. We treat an LLM-backed feature as a system that has to be tested, measured and observed, exactly like any other piece of production. You don't get another demo — you get a feature where it's clear what gets checked, when, and with what result. Sometimes the analysis ends with us saying you don't need a model at all: plenty of what companies ask for "with AI" is really a data or a process problem, cheaper and more predictable to solve with ordinary automation.
An LLM-backed feature reaches customers only when you can test it, observe how it behaves and account for what it costs. That is how model work is done in environments under regulatory supervision — PwC and Roche among them, where the firm's founder worked. In a smaller company those three things cost barely more, and they decide whether the feature ships at all.
What the service covers
Discovery and feasibility
You leave the workshop knowing what to invest in and what to leave alone. We start by walking through the processes where AI could change something and by assessing each one soberly: whether a model is needed or a script would do, whether the data is usable at all, what it will cost per month and what happens when the model gets it wrong. The result is a short, prioritised list — including the items we recommend against.
- Review of processes and a shortlist of candidates
- Data readiness: quality, availability, access rights
- The call: language model or plain automation
- Cost and risk estimate before any work starts
Building LLM-backed features
The result lives in the tools your team already works in — not on a separate "AI platform". We build specific features: an assistant that answers from your own documents, search across internal knowledge, structured extraction from invoices, contracts and email, ticket classification, drafts prepared for a person to approve.
- Chatbots and assistants answering from your own documents
- Retrieval over internal knowledge (RAG) instead of digging through folders
- Structured extraction from documents and email
- Ticket classification and drafted replies
- Integration with existing systems and APIs
AI quality and evaluation
This is the part that usually decides whether a feature survives contact with customers. We build an evaluation set from real cases in your company and run it on every prompt change and every new model version. Out-of-scope questions get checked separately, along with attempts to pull an answer the model doesn't actually have. Acceptance criteria are agreed before anything reaches a customer — you hear about weak answers before your customers do, not from them.
- An evaluation set built from real cases in your company
- Regression tests for prompts and successive model versions
- Checks for invented answers and out-of-scope questions
- A human review loop wherever the stakes are high
- Acceptance criteria agreed before the feature ships
Observability and cost control
In production you need to see what the model is actually doing. We wire up tracing for every call — the question, the context, the answer, the latency, the tokens burned — in an LLM observability tool; Langfuse is one we work with. On top of that, alerts for when quality starts sliding or cost drifts away from the plan.
- Tracing of model calls: question, context, answer, latency
- Token cost monitoring and limits that protect the budget
- Alerts when answer quality starts to degrade
- Collecting user feedback and folding it back into the evaluation set
Security, data and compliance
We settle explicitly what data may go to a model and what must never leave the company, then design to that decision. We review the implementation for the failure modes specific to LLMs: instructions injected through document content, data leaking into answers, API keys sitting in client-side code. This is secure design and review, not formal penetration testing.
- Deciding what data may reach a model and what may not
- Review for prompt injection and data leakage
- Keys and model calls on the server, never in the browser
- Retention of logs and conversations, GDPR and EU processing questions
Typical situations
A pilot chatbot went down brilliantly in the demo, and now nobody in the company will sign off on letting customers near it.
The support inbox is drowning in questions whose answer hasn't changed in three years.
Company knowledge sits in PDFs, procedures and SharePoint folders, and finding anything means asking one specific person.
The board is asking "what are we doing about AI", and there's no shortlist of concrete use cases and no estimate of what they would cost.
What you can count on
- A working feature wired into the systems people use daily, not a prototype living on a separate link.
- A ship-or-hold call backed by evidence — the evaluation suite runs on every change to a prompt, a model version or a knowledge source, with a result you can put in front of a room.
- You see what the model really tells users and what it costs to run — call tracing and a cost view instead of guesswork.
- A written recommendation, including what is not worth handing to a model and why.
Technologies
- Python
- TypeScript
- LLM APIs
- RAG
- Langfuse
- PyTest
- PostgreSQL
- AWS
- CI/CD
Projects in this area
Selected work where this scope was part of the delivery.
ThomasSeekerAI
A private trader was combing OLX, Otomoto, Allegro and a handful of smaller sites by hand, every day, for things worth reselling. We built an internal tool that takes that work over: searches described in a plain sentence run on a schedule in the background, the noise is filtered out, and what is left is ranked by how good the opportunity looks — hours of manual browsing turned into an automated job.
Private opportunity-discovery tool with AI-assisted analysis
Questions about this service
Not always, and it's the first thing we settle, because it determines the architecture. Some use cases work fine with processing in an EU region and training on submitted content switched off. Some data can be anonymised or simply never sent. Where requirements are strict, we consider a model running on your own infrastructure — with the honest caveat that it costs more and usually performs worse. We explain what each choice buys you rather than assuring you that "it's all secure".
Because we measure it. We collect real questions and documents from your company along with the answers you expect, then run that set on every change: prompt, model version, knowledge sources. You see not just that things feel better, but which cases improved and which broke. Out-of-scope behaviour gets checked separately, because a well-built assistant is supposed to say it doesn't know rather than invent something that sounds credible.
Plain automation is often cheaper, and we say so in the first conversation. A language model earns its keep where free-form text is involved: questions asked a hundred different ways, documents with inconsistent layouts, content no fixed rule can describe. If the task has clear rules, a script will be cheaper, faster and predictable. We would rather lose an AI project than sell a model where fixing the process would have been enough.
We finish by handing you the tools: call tracing, a cost view, alerts and an evaluation suite your team can run on its own. Models and their versions do change on the provider's side, though, so for customer-facing features we offer ongoing care: regular reviews of answer quality and a response when something starts slipping. We work remotely with companies across Poland, and on site in and around Bydgoszcz — a kickoff workshop goes better around one table.
Related services
QA and software testing
Manual and automated testing, API, UI and regression tests, plus test strategy — a quality process that actually protects your releases.
Automation and integrations
We replace copying data between systems, manual reports and retyping documents with scripts and integrations that run on their own.
Software development
Software written for a specific problem: internal systems, MVPs, integrations and taking over an application from another vendor — designed around how your company works, not the other way round.
Let's check whether AI solves a real problem for you
The first conversation is free. Tell us what AI is supposed to do in your company and you'll get a concrete answer: what we'd do — a language model or plain automation — how we'd measure that it works, what it may cost, and where to start. If a model wouldn't pay for itself here, we'll say so straight out.