Course 9, lesson 88 of 100, Ages 14+
Speed, cost and tokens
Running AI efficiently
Like I’m 5
Every time an app asks an AI something, it costs a little money and time. Smart builders make AI apps fast and affordable.
The big idea
AI APIs usually charge per token, for both the text you send and the text you get back. Long prompts, huge documents and long answers cost more and take longer.
To save, trim prompts, cache repeated answers, use smaller models for easy tasks and bigger ones only when needed, and stream replies so users see text appear quickly. Measure latency and cost per request just like you measure accuracy.
Examples
- Right-sizing: A small model sorts emails; a large one writes tricky replies.
- Caching: Common questions reuse a stored answer.
- Streaming: Words appear as they're generated, so the app feels fast.
How it works
- Measure tokens, cost and response time per request.
- Trim prompts and pick the smallest model that does the job well.
- Cache repeats and stream responses.
Check your understanding
- What do AI APIs usually charge for?
- Options: Tokens in and out; The number of screens; Each keyboard press.
Answer: Tokens in and out. Input and output tokens determine the cost. - Why use a smaller model for easy tasks?
- Options: It's cheaper and faster when it's good enough; It's always smarter; Big models can't do easy tasks.
Answer: It's cheaper and faster when it's good enough. Match the model to the task's difficulty.
Remember
Cost and speed depend on tokens and model size. Measure, trim, cache and right-size.
Talk about it
Which part of an AI app would you cache? Why?
Go deeper
Techniques include prompt caching, batching, model routing, distillation and quantisation. Time-to-first-token and tokens per second are key latency metrics.