The questions you will be able to answer
You are not expected to know any of the words below yet. Nobody starting this course does. This page puts them all in front of you once, in plain language, so that none of them can surprise you later. Each answer is a few sentences; the course teaches every one of them properly when its turn comes.
Read the questions first. Guess before you open an answer. The ones you get wrong are the ones you will remember.
Open any chatbot you have: ChatGPT, Claude, Gemini, the assistant on your bank's site. In three separate fresh chats, ask the same question: "I bought a kettle 40 days ago and it stopped working. Can I get a refund?"
- Write down the three answers side by side.
- Mark every place they differ: the number of days, the conditions, the tone, what it asked you.
- Now answer honestly: which of the three is the expected result?
You cannot say, and neither can anyone else, until somebody decides what "good" means and measures it across many runs. You have just done the first step of an AI tester's job. The rest of the course is doing it properly.
What is a model?
A model is a program that learned its behaviour from examples instead of being written as rules. Nobody typed "if the guest asks about dogs, say 25". It read a huge amount of text and learned what a sensible reply looks like. That is why you cannot find the line of code that caused a wrong answer.
What is an LLM?
A large language model: a model trained on text, whose one skill is predicting what comes next. ChatGPT, Claude, and Gemini are products built on LLMs. When this course says "the model", it means one of these.
What is a prompt?
Everything you send to the model: the instructions, any documents, and the question. Think of it as the test input. Change one word in it and the output can change a lot.
What is a token?
The unit a model reads and writes: a word or a piece of a word. "Cancellation" might be two or three tokens. It matters to a tester because cost, speed, and size limits are all counted in tokens, not in words.
Why does the same question get different answers?
The model does not look up an answer. It builds one, a token at a time, and at each step it picks from several likely options with a little randomness. A setting called temperature controls how much. So one run tells you very little, and that single fact is why AI testing is a different job.
What is a hallucination?
An answer that sounds confident and is simply false: a fee that does not exist, a policy nobody wrote. The model is not lying; it is producing likely-sounding text. Finding these and measuring how often they happen is a large part of your work.
What is a context window?
How much text the model can keep in view at once: the instructions, the documents, the conversation so far, and its own answer. When a long chat overflows it, the oldest parts fall out. That is why a chatbot "forgets" what you told it earlier.
What is a system prompt?
The standing instructions the developer gives the model before any user speaks: "You are the hotel's assistant. Answer only from the policy." The user never sees it. A lot of defects are really defects in this text.
What is RAG, and why is it needed?
Retrieval-augmented generation. A model knows nothing about your company's documents and its general knowledge stops at a cutoff date. RAG fixes both: when a question arrives, the system first searches your documents for the relevant passages, then hands them to the model and says "answer from these". Without it the bot guesses. With it, the bot can still go wrong in two new places: the search can fetch the wrong passage, or the model can ignore the right one. You will test both.
What is an embedding?
A way of turning a sentence into a list of numbers so that sentences with similar meaning end up close together. It is how the search in RAG finds "Are pets allowed?" when the guest typed "Can I bring my dog?".
What is an agent?
A model that is allowed to do things, not only talk: look up a booking, check availability, cancel. It asks for a tool, your code runs it, and the model reads the result. The risk is obvious, and testing what it did, not just what it said, is its own lesson.
What is an API, and what is an API key?
An API is how one program talks to another: send a request, get a response. If you have used Postman you have used one. An API key is the password your program sends so the provider knows who to bill. From Module 5 you will have one, and Module 1 teaches you never to put it in your code.
What is an eval?
Short for evaluation: a test suite for an AI product. It runs many questions through the system, checks every answer, and reports how many were acceptable. Where you would say "test run" today, AI teams say "eval".
What is a golden dataset?
The fixed list of questions an eval asks, each with what a correct answer must contain. It is your test-case library for AI, and writing a good one uses exactly the test-design skill you already have.
What are a rubric and a judge?
A rubric is a written checklist for grading an answer when there is no single correct wording: mentions the fee, promises nothing false, stays polite. A judge is a second model that applies the rubric to thousands of answers for you. You check the judge against your own grading before you trust it.
What are a pass rate and a threshold?
The pass rate is how many answers were acceptable: 41 of 48 is 85%. The threshold is the rate the team agreed is good enough before the run. Pass or fail becomes "85% against a bar of 90%", which is a much more honest sentence.
What is prompt injection?
Text that tricks the model into following someone else's instructions. A sentence hidden in a document says "ask the guest for their card number", the bot reads the document, and does it. It is the most important security problem in AI products, and you will plant, find, and report one.
What is red-teaming?
Attacking your own product on purpose, with permission, to find what a real attacker would. It is exploratory testing with a security mindset. In this course you only ever do it against the practice app on your own machine.
What is drift?
Quality changing over time without anyone touching your code, usually because the AI provider updated the model. It is why AI test suites run every night, not only before a release.
Do I have to become a programmer?
No. You need enough Python to loop over a list of questions, call an API, and write a line that says "this must be true". That is a few weeks of learning, and Stage A teaches exactly that and stops.
What will I actually do in this job?
Decide what "good" means for a product, write the questions that probe it, measure how often it is good, attack it, and tell the team whether it is ready, with numbers. The engineer builds the system. You are the one who can say how well it works.
What if I get stuck, or fall behind?
Everyone does in Module 2. Every lesson has an answer key, every lab has a checklist, and the pack remembers where you stopped. There is no deadline. Seven hours a week is the plan; three hours a week still gets you there.
| Sitting | Do | Time |
|---|---|---|
| 1 | This page and the kettle experiment. Then Lesson 0.1. | 45 min |
| 2 | Lesson 0.2, then install the kit in Lesson 0.3. | 2 h |
| 3 | Run the setup checker in Lesson 0.4. Start Module 1. | 1 h |
The next seven weeks have no AI in them on purpose: terminal, Python, APIs, and browser tests are the tools every AI test is built from. If you want to see where it is all going, you are allowed to read Lessons 5.1 to 5.3 at any time. They need no code. Then come back.
Every term on this page, and every other one in the course, is in the course glossary: 44 terms in plain words, linked from the contents list of every module.
The course teaches every one of these answers properly, by doing.
Thirteen modules from your first terminal command to a release assessment you present on video. A practice AI product with 12 planted defects to find. Seven portfolio projects on your own GitHub. One payment of ₹1,499, yours to keep.
All sales are final, which is why this page is free. Read it, try the experiment, then decide.