A "ChatGPT application" in 2026 is usually a product that calls a large language model through an API, grounds it in your own data, lets it use a few tools, and wraps all of that in tests, guardrails and monitoring. The model call is the easy part; the surrounding system is where the time and money go.
This guide is for founders and product owners scoping an LLM-powered app, and for engineers who want a checklist. It covers what you are actually building, the architecture choices that matter, realistic effort drivers, and how to judge whoever you hire to build it.
Three things people mean by "ChatGPT app"
Before you scope anything, decide which of these you need. They have very different costs and constraints.
- A custom GPT inside ChatGPT. A configured assistant with instructions, uploaded files and optionally a few API actions. Cheap to make, lives entirely inside ChatGPT, and gives you little control over UX, analytics or data flow. Good for internal helpers and prototypes.
- An app that runs inside ChatGPT. OpenAI's Apps SDK (built on the open Model Context Protocol) lets you expose your service's tools and UI components to ChatGPT users in the conversation. You build an MCP server and widgets; ChatGPT handles the chat. This is a distribution play: you reach ChatGPT's users, but on its terms.
- Your own application that calls a model API. Your web or mobile app, your backend, your data, with an LLM as one component. This is what most "ChatGPT application development" projects really are, and most of this guide is about it. The same architecture works with OpenAI's models or with any comparable provider.
The architecture of a production LLM app
Strip away the branding and nearly every serious LLM application has the same layers.
| Layer | What it does | Typical decisions |
|---|---|---|
| Client | Chat or task UI, streaming output, citations, feedback buttons | Streaming vs. full response; how to show sources; undo for actions |
| Orchestration backend | Builds the prompt, calls the model, runs tool calls, enforces limits | Plain code vs. a framework; sync vs. queued jobs |
| Model provider | Generates text, structured output and tool calls | Which model per task; fallback provider; data-retention terms |
| Retrieval (RAG) | Finds relevant chunks of your documents or records | Chunking, embeddings, hybrid keyword+vector search, reranking |
| Tools | Functions the model may call: search, CRM lookups, ticket creation | Read-only vs. write access; confirmation steps; per-user permissions |
| Guardrails | Input/output checks, PII handling, policy filters | What to block, redact or escalate to a human |
| Evals & observability | Test sets, traces, cost and latency dashboards, user feedback | What "good" means; who reviews failures; regression gates |
Retrieval-augmented generation, done properly
RAG means fetching relevant passages from your own content and putting them in the prompt so the model answers from them rather than from memory. It is the default way to make a model "know" your product docs, policies or knowledge base, and it is far cheaper to update than fine-tuning.
The quality of a RAG system is decided mostly outside the model:
- Ingestion. Clean extraction from PDFs, HTML and tables. Garbage parsing produces confident garbage answers.
- Chunking. Split on document structure (headings, sections), keep titles and metadata with each chunk, and avoid splitting tables mid-row.
- Search. Combining vector similarity with keyword search (BM25 or similar) catches exact terms like product codes that embeddings miss. A reranking step on the top results usually improves precision.
- Access control. Filter retrieved chunks by what the current user is allowed to see. Skipping this is how an internal assistant leaks HR documents.
- Citations. Ask the model to cite chunk IDs, then render them as links. Users trust answers they can check, and you can measure whether citations actually support the claim.
Long context windows have not made retrieval obsolete. Stuffing everything into the prompt costs more per call, adds latency, and models still miss details buried in very long inputs.
Tool use and agents
Modern model APIs let you describe functions with a JSON schema; the model decides when to call them and with what arguments, your code executes the call, and the result goes back to the model. That is how an assistant checks an order status or books a meeting.
Treat every tool as an API exposed to an untrusted caller. The model can be steered by text inside documents or emails it reads (prompt injection), so:
- Run tools with the end user's permissions, never a superuser key.
- Require explicit confirmation for anything that spends money, sends messages or deletes data.
- Validate arguments server-side as you would any form input.
- Cap the number of tool-call steps per request so a confused agent loop cannot run up a bill.
Multi-step "agents" are worth it when the task genuinely branches. For a fixed sequence (classify, retrieve, draft, check), a plain pipeline is cheaper, faster and easier to debug. The same trade-offs show up in on-chain automation; see this site's piece on AI agents and autonomous smart contracts.
Structured output
If the model's output feeds code, request JSON that conforms to a schema; the major APIs support schema-constrained output. Then validate it anyway. Free-text parsing with regular expressions is a common source of brittle production bugs.
Evals: how you know it works
An LLM feature without an evaluation set is a demo. Before launch you want a few hundred representative inputs with expected behavior, drawn from real user questions where possible, including the awkward ones: out-of-scope requests, ambiguous wording, adversarial prompts, questions whose answer is "I don't know".
- Deterministic checks for anything checkable: valid JSON, correct tool chosen, required fields present, no banned phrases.
- Reference comparisons for retrieval: did the right document appear in the top results?
- Model-graded checks for fuzzy qualities like faithfulness to sources or tone, calibrated against human ratings on a sample so you know the grader agrees with people.
- Human review of a slice of production traffic every week, feeding new cases back into the test set.
Run the suite on every prompt change and every model version change. Providers retire and update models regularly; an eval set is what lets you switch with confidence instead of hoping. Prompt craft and eval design are covered in more depth in the guide to prompt engineering.
Cost and latency: where the money goes
Model usage is billed per token, with output tokens typically priced higher than input tokens. The per-request cost is roughly:
cost ≈ (input_tokens × input_price) + (output_tokens × output_price)
input_tokens = system prompt + conversation history + retrieved chunks + tool results
Prices change often, so model your costs from the provider's current price page rather than any figure in an article. The levers are stable, though:
- Retrieved context is usually the biggest input. Retrieving eight good chunks instead of thirty mediocre ones cuts cost and often improves answers.
- Prompt caching. Providers discount repeated prompt prefixes. Put the stable parts (instructions, tool definitions) first and the variable parts last.
- Model routing. Use a small, fast model for classification and simple answers; reserve the largest or reasoning-heavy models for the hard cases.
- Conversation trimming. Summarize or drop old turns rather than resending an ever-growing history.
- Batch APIs for work that does not need to be interactive, such as nightly document tagging, usually at a discount.
For latency, users care most about time to first token. Stream responses, run retrieval and other lookups in parallel, and avoid chaining several sequential model calls on the hot path. Reasoning models think before they answer, which helps on hard problems and hurts on simple ones.
Privacy, security and compliance
- Data terms. Read the provider's API data-usage and retention terms. Business API traffic is generally not used for training by the major providers, and some offer zero-data-retention arrangements for eligible customers. Confirm in writing rather than assuming.
- Data residency. If you serve EU users or regulated industries, check where requests are processed and whether a regional endpoint or a cloud-hosted deployment of the model is available.
- PII minimization. Do not send what the model does not need. Redact identifiers before the call where you can.
- Logging. Your traces contain user inputs. Apply the same retention and access rules you apply to the rest of your user data.
- Regulation. The EU AI Act imposes transparency duties (for example, telling people they are talking to an AI) and heavier obligations for high-risk uses such as hiring or credit decisions. If your use case touches those areas, get legal input early.
The build process
- Pick one job to be done. "Answer support questions about billing using our help center" beats "an AI assistant for customers". Define what a good answer looks like and when it should hand off to a human.
- Collect real examples. Pull 100–300 actual questions or tasks from tickets, search logs or interviews. This becomes your first eval set.
- Build a thin prototype. One prompt, basic retrieval, no polish. Run the eval set. You will learn more from the failures than from any planning document.
- Fix retrieval and data first. Most early failures trace back to missing, stale or badly chunked content, not the model.
- Add tools and guardrails. Introduce actions one at a time, each with permission checks, confirmation flows and tests.
- Harden for production. Rate limits, retries with backoff, provider fallback, cost caps per user, tracing, dashboards.
- Launch to a small group. Watch traces daily, label failures, extend the eval set, and only then widen access.
- Operate it. Budget ongoing time for content updates, model migrations and reviewing flagged conversations. LLM features are never "done".
If you are still testing whether the idea is worth funding, the same discipline applies at smaller scale; the guide to proof-of-concept development covers how to scope a POC so it answers a real question.
Realistic effort and cost drivers
Rather than quote prices that vary by region and seniority, here are reasoned effort ranges. Multiply by your team's loaded weekly rate.
| Scope | Typical team | Rough duration |
|---|---|---|
| Internal Q&A assistant over one document set, web UI, basic evals | 1–2 engineers | 3–6 weeks |
| Customer-facing support assistant with RAG, handoff to human agents, analytics | 2–3 engineers + designer part-time | 8–14 weeks |
| Agentic workflow writing to business systems (CRM, ERP, ticketing) with approvals | 3–4 engineers + QA | 3–6 months |
The ranges widen with messy source data, strict compliance requirements, multiple languages, and every integration with a legacy system. Ongoing model usage is a separate line item that scales with traffic and context size. Hooking an LLM into existing tools rather than building a new product is a different, usually smaller, project; see ChatGPT integration for that path.
How to evaluate a team or vendor
Ask questions that reveal whether they have shipped and operated LLM systems, not just demoed them:
- How will you measure answer quality before and after launch? Show me an eval set from a past project (anonymized).
- How do you handle prompt injection through retrieved documents or tool results?
- What happens when the provider deprecates the model we launch on?
- How do you enforce per-user document permissions in retrieval?
- What is your estimate of cost per conversation, and what assumptions is it based on?
- Will we own the prompts, eval data, and code, and can we switch model providers without a rewrite?
Red flags: guarantees of a specific accuracy percentage before seeing your data, fine-tuning proposed as the first step, no mention of evaluation, or an architecture that locks you to one proprietary platform for no stated reason.
For broader AI work beyond language models, such as forecasting or computer vision, the guide to choosing an AI development partner covers different evaluation criteria.
When not to build an LLM app
If the task has a correct answer computable by rules, a database query or a form, build that instead. LLMs are a good fit for language-heavy, variable inputs where a mostly-right draft saves real time and a human or a check catches the rest. They are a poor fit where every error is costly and nobody reviews the output.
Frequently asked questions
Do I need to fine-tune a model for my app?
Usually not at first. Good instructions plus retrieval over your own data solve most "it doesn't know our business" problems, and they are easier to update. Fine-tuning helps with consistent format or style at scale, or with narrow classification tasks, once you have an eval set that shows prompting alone falls short.
Is it risky to depend on one model provider?
Somewhat. Keep provider-specific code behind a thin interface, keep your prompts and evals in your own repository, and test at least one alternative model against your eval set. That makes switching a measured decision rather than an emergency.
How do I stop the assistant from making things up?
Ground it in retrieved sources, instruct it to answer only from them and to say when it cannot, require citations, and measure faithfulness in your evals. You can reduce fabricated answers a great deal this way, but not to zero, so design the UX for verification on anything high-stakes.
Can I use an open-weight model instead of a hosted API?
Yes. Self-hosting an open-weight model gives you full control over data and can be cheaper at high, steady volume, but you take on GPU capacity, serving infrastructure and upgrades. Many teams start with a hosted API and revisit once usage and requirements are clear.
How long until an LLM app is production-ready?
A narrow internal tool can be useful in a few weeks. A customer-facing assistant with integrations, evals and monitoring more commonly takes two to four months, mostly spent on data quality, edge cases and operational hardening rather than on the model.