Sound familiar?
- A prototype built in a notebook or no-code tool cannot be put in front of real users
- Outputs vary from run to run and nobody can say whether a change made things better
- Model costs are higher than expected and nobody knows which feature is responsible
- You are locked into one provider's SDK throughout your codebase
- There is no audit trail of what the model was asked and what it said
- The app slows to a crawl whenever the model provider is busy
Key facts
- We build the whole application, not just the prompt: users, roles, database, queues and admin screens
- Structured output validated against a JSON schema before it is saved or shown
- An automated evaluation suite runs on every prompt or model change
- Provider-neutral service layer so you can change model without a rewrite
- Streaming responses, background queues and retries for slow or failed calls
- Cost per task tracked, with prompt caching and model routing to keep it down
- You own the code, prompts and evaluation data
From prototype to product
Getting a language model to produce a good answer in a playground takes an afternoon. Turning that into software your team or customers use every day takes engineering. An LLM application still needs everything a normal web application needs, such as login, permissions, a database, background jobs and backups, plus a few things that are specific to models: evaluation, output validation, cost tracking and a plan for when the provider is slow or down.
We build in the stacks we use for all our work: PHP 8 and Laravel, Node.js and TypeScript with Next.js, or Python with Django or FastAPI when the project leans on Python's data libraries. The model is one component in the architecture, not the whole thing.
Typical LLM applications we build
The projects that work best take a task your team already does by hand, involving a lot of reading or writing, and put a model in the middle with a person still in charge. Examples:
- Quote and proposal drafting: the app pulls product data, pricing rules and past wording, drafts a document, and a salesperson edits and approves it.
- Ticket and enquiry triage: incoming messages are categorised, prioritised and summarised, with a suggested reply waiting for the agent.
- Report writing: inspection notes, site visit photo captions or survey answers become a first-draft report in your house style.
- Data clean-up: thousands of free-text product descriptions or addresses are normalised into structured fields, with low-confidence rows flagged for review.
Each one is a normal application with an LLM step, which keeps it testable and keeps your team in control of what goes out.
Structured output and validation
Free text is hard to build software on. Wherever we can, we ask the model to return JSON that matches a schema, using the structured output or tool calling features that OpenAI, Anthropic and Google all now offer. Our code then validates the response, for example with Zod in TypeScript, Pydantic in Python or Laravel validation rules in PHP. If it fails, we retry with the error message or route it to a person. Nothing unvalidated is saved to your database or rendered as HTML.
This also makes prompts testable. A prompt that classifies support tickets either returns one of your twelve categories or it does not.
Evaluation: how you know a change helped
Changing a prompt or upgrading a model can quietly make some answers better and others worse. We treat that like any other regression risk. Each application gets an evaluation harness:
- A set of real inputs with expected outputs, agreed with your team
- Automatic scoring: exact match for fields and categories, similarity or rubric-based scoring for written text
- Where useful, a second model used as a grader with a written rubric, spot-checked by a person
- A report comparing the new version to the current one on accuracy, latency and cost
The harness runs in the deployment pipeline, so a change that drops accuracy below the agreed threshold does not go live.
Speed, reliability and cost
Model calls are slow compared with a database query, and providers have rate limits and occasional outages. We design for that:
- Streaming responses so users see text arrive rather than waiting on a spinner
- Queues (Laravel Horizon, BullMQ or Celery, usually backed by Redis) for bulk and background work
- Timeouts, retries with exponential backoff and a fallback model where appropriate
- Model routing: a small, fast model for simple tasks and a larger one only where it is needed
- Prompt caching for long, repeated instructions or documents where the provider supports it
Every call is logged with its token counts, so the dashboard shows cost per feature and per user, not just one monthly bill.
Security and data
LLM applications bring new risks alongside the usual ones. We follow the OWASP Top 10 for LLM Applications, which in its 2025 edition puts prompt injection first and also covers sensitive information disclosure, improper output handling and unbounded consumption. In practice that means least-privilege tools, output escaping, per-user rate limits and spend caps, and keeping secrets and other customers' data out of prompts entirely.
For personal data we document what goes to which provider, in which region and under what retention terms, so your DPIA is straightforward. If data cannot leave your infrastructure, we can design around an open-weight model on your own servers or private cloud. For broader context see AI integration, and to discuss a build, get in touch.
What we deliver
- A web application in Laravel, Next.js or Django with authentication and role-based access
- A model service layer with provider adapters, retries, timeouts and fallbacks
- Prompt templates kept in version control with typed inputs and schema-validated outputs
- An evaluation harness with test cases, scoring and a regression report
- Background job processing for long-running or bulk tasks
- Logging and dashboards for latency, error rate, accuracy and cost per task
- Deployment, documentation and handover
How it works and what it costs
Every project gets a fixed-price quote after a free initial chat and a short scoping stage. You own the code and the data.
Free chat
Tell us the problem in plain English: what you do now, what goes wrong and what "better" looks like. No charge, no obligation.
Scoping
We map the processes, systems and data involved, agree what is in and out, and write it down so there are no surprises.
Fixed-price quote
You get a fixed price for the agreed scope, or a phased plan for bigger builds, so you can start small and prove it works.
Build and test
We build in short stages you can see and try, test against real data, then go live carefully with a rollback plan.
Hand over and look after
You own the code and the data. We can host it, support it and keep improving it, or hand it to your own team.
Frequently asked questions
What is an LLM application?
An LLM application is software where a large language model does part of the work, such as drafting, classifying, extracting or summarising text, inside a normal application with users, permissions and data storage. dijitul developments builds the whole application around the model, including evaluation, validation, logging and cost control.
Can you take over our existing AI prototype?
Yes. We review the prototype, keep the prompts and examples that work, and rebuild it as a maintainable application with tests, an evaluation suite and proper deployment. Often the prompts are a good starting point and the missing parts are validation, error handling and security.
Which programming languages do you use?
Mostly PHP 8 with Laravel, Node.js and TypeScript with Next.js, and Python with Django or FastAPI. We choose based on your existing systems and team. All the main model providers offer official SDKs or HTTP APIs that work with these stacks.
How do you stop us being locked into one AI provider?
We put the model behind a thin service layer in your code, with adapters for each provider. Prompts, schemas and the evaluation set are provider-neutral. To switch, we add or change an adapter, run the evaluation suite and compare results, rather than rewriting the application.
Can an LLM app run without sending data to the cloud?
Yes. Open-weight models can run on your own servers or a private cloud with GPU capacity. They are usually less capable than the largest hosted models, so dijitul developments tests them against your evaluation set first to check they meet the accuracy you need.
How is an LLM app project priced?
dijitul developments quotes a fixed price for an agreed scope after a free chat and a discovery stage. Larger applications can be split into phases, each with its own fixed price. Model usage is billed separately by the provider, and we estimate it in advance.
Related
Tell us what you need to build
Free chat, clear scope, fixed-price quote. You own everything we build.