# AI Mastery Lab

30 practical lessons · DiscoveryVIP

This tool contains authored lessons and deterministic local experiments. It does not call an AI model. Practice uses fictional data.

## 1. Choose a job for AI

AI is a broad family of systems that recognize patterns, make predictions or generate content. Start with a task and an observable outcome, not a tool name. A useful pilot has clear inputs, a manageable cost of error and a person who can judge the result.

**Example:** Drafting three alternative meeting agendas is easier to review than delegating a final hiring decision.

**Try it:** Name one repetitive task, its input and what a good result looks like.

**Check:** Which is the best first pilot?

**Answer:** Draft an agenda that you can review — A draft has a clear purpose and a review step before anyone relies on it.

## 2. Rules, prediction or generation?

A fixed rule follows a condition written by a person. A predictive model estimates a label or value from learned patterns. A generative model produces new content. Many products combine all three. Use a simpler rule when the requirement is exact and stable.

**Example:** A rule can flag an invoice over $500; a classifier can estimate its category; a generator can draft a follow-up.

**Try it:** Break a workflow into one rule, one prediction and one generation step.

**Check:** Which task can a simple deterministic rule handle?

**Answer:** Flag a value greater than 500 — A numeric threshold is explicit; it does not need a learned model.

## 3. Training versus using a model

Training adjusts model parameters using data and an objective. Inference uses a trained model to produce an output for an input. Giving context in a conversation is not the same operation as retraining model weights. Product data policies determine what providers retain or use later.

**Example:** Pasting a style guide can affect the next answer without permanently teaching that style to the underlying model.

**Try it:** Describe what context your task needs only for the current request.

**Check:** Does adding a paragraph to a prompt necessarily retrain the model?

**Answer:** No, it supplies input context — Context can guide a response without changing the model parameters.

## 4. What language models produce

Language models generate token sequences using learned patterns and the current context. Their fluent output can be useful without being a verified record of reality. A product may also retrieve documents or call tools, so inspect what evidence it actually used.

**Example:** A convincing project summary can still invent an owner when the notes never named one.

**Try it:** Write an instruction for handling a missing owner.

**Check:** Which is evidence that an owner was assigned?

**Answer:** An explicit statement in the source notes — Plausibility is not a substitute for a source statement.

## 5. Context has limits

A context window holds the material a model can process for a request, subject to model and product limits. Long input can crowd out useful details or make retrieval difficult. Organize the relevant material and avoid assuming a previous message will always be available.

**Example:** For a policy question, include the applicable section and version rather than unrelated reports.

**Try it:** List the minimum context needed for a policy answer.

**Check:** What is the better context strategy?

**Answer:** Provide the relevant section with its date — Relevant, labeled material makes the intended evidence easier to identify and inspect.

## 6. A score is not certainty

A classifier score is a model output used to rank or select decisions. Whether it represents a reliable probability depends on calibration and validation. A language model’s verbal confidence is different again. Decide how errors will be handled before choosing a cutoff.

**Example:** A support classifier with score 0.8 can still be wrong. A lower threshold may catch more urgent tickets but create more false alarms.

**Try it:** Move the threshold and describe one changed decision.

**Check:** Does a high score prove a prediction is correct?

**Answer:** No, compare it with labeled outcomes — Scores inform decisions but require evaluation against actual outcomes.

## 7. Define a useful deliverable

Specify what you need produced and how you will use it. “Help with this” leaves the output ambiguous. A concrete deliverable gives you something testable: a summary with named fields, a set of alternatives or a draft message for a particular reader.

**Example:** Ask for three agenda items, each with a purpose and a time allocation.

**Try it:** Build a prompt that names the deliverable and intended use.

**Check:** Which instruction is easiest to evaluate?

**Answer:** Return three agenda items with minutes — Explicit fields and counts can be checked against the output.

## 8. Name the audience

Audience affects vocabulary, assumed knowledge, tone and level of detail. Describe the reader’s needs directly. A role label such as “expert” does not grant expertise or ensure correctness; the output still needs a review matched to the task.

**Example:** An explanation for a first-time spreadsheet user should define “range” before using it repeatedly.

**Try it:** Rewrite a technical request for a beginner audience.

**Check:** What most directly improves audience fit?

**Answer:** Describe the reader’s knowledge and needs — Audience information guides language and explanation depth.

## 9. Supply the facts that matter

Separate task instructions from source material. Identify which facts must be preserved and which details are unknown. Avoid silently asking the model to fill missing facts with guesses. If an answer depends on absent information, request a question or an explicit unknown.

**Example:** Notes say the workshop lasts 90 minutes but do not name a room. The room remains unknown.

**Try it:** Create a brief that identifies an important missing fact.

**Check:** What should happen to an absent room number?

**Answer:** Mark it unknown or ask for it — Missing information should stay visibly missing until verified.

## 10. Constrain the format

A format requirement makes the output easier to use. Name the fields and give a small example if needed. Requesting JSON alone does not prove the result is valid or complete; software should parse and validate it before use.

**Example:** For action items, request task, owner and due_date; use null for missing values.

**Try it:** Define three fields and how missing information should appear.

**Check:** What should software do before using generated JSON?

**Answer:** Parse and validate it — Syntax and field validation catch different kinds of errors.

## 11. Show a representative example

Examples can demonstrate a style, label boundary or output shape more efficiently than a long description. Include realistic edge cases, not just perfect inputs. A model may imitate mistakes or irrelevant details in the example, so keep it intentional.

**Example:** Show one action with a known owner and another with owner: null.

**Try it:** Write a short input/output example for your task.

**Check:** What makes an example especially useful?

**Answer:** It demonstrates the intended boundary or format — An example should teach the behavior you actually want repeated.

## 12. Iterate with a controlled change

When output disappoints, identify the failure and change one meaningful part of the request. Keep the test input stable so you can compare results. A longer prompt is not automatically better; remove instructions that conflict or distract.

**Example:** If owners are invented, add a rule to use null when the source does not assign one, then rerun the same notes.

**Try it:** Write a before/after change and one check for success.

**Check:** How can you tell whether a prompt change helped?

**Answer:** Compare it on the same test cases — Controlled comparison makes it easier to connect a change with its effect.

## 13. Break answers into claims

An answer may mix supported facts, unsupported additions and contradictions. Review one claim at a time. This is more informative than deciding whether the entire paragraph “sounds right.” A claim about a quantity, date or permission deserves its own check.

**Example:** A source can support a 90-minute duration while saying nothing about a completion certificate.

**Try it:** Classify the claims in the evidence desk.

**Check:** What should you review independently?

**Answer:** Each meaningful factual claim — Different claims in the same answer can have different evidence quality.

## 14. Supported, contradicted or unknown?

Supported means the source provides the needed evidence. Contradicted means it states something incompatible. Unknown means it does not settle the claim. Absence of evidence is not automatically a contradiction. Keep those three categories separate.

**Example:** A note that gives no room does not prove the workshop is online.

**Try it:** Find a claim that the source leaves unknown.

**Check:** The source never mentions a certificate. What is “A certificate is included”?

**Answer:** Not established by this source — A missing detail remains unknown unless another reliable source resolves it.

## 15. Check citations at the source

A citation is useful only when the source exists, is relevant and supports the specific claim. Open it and inspect the passage, date and scope. A real document can be cited for a statement it never made.

**Example:** A 2022 policy may be authentic but superseded by a later version.

**Try it:** Write three checks you would perform on a citation.

**Check:** What is the strongest citation check?

**Answer:** Read the relevant passage and its scope — Citation quality depends on support, not appearance.

## 16. Understand retrieval

Retrieval selects potentially useful material before an answer is produced. It may use keywords, embeddings or a combination. Finding related text is not the same as finding decisive evidence. The local experiment uses transparent word overlap, not a trained embedding model.

**Example:** A document mentioning workshop certificates may still refer to a different course.

**Try it:** Search the evidence desk and inspect why each document matched.

**Check:** Does a top-ranked document necessarily prove a claim?

**Answer:** No, relevance and support are different — A retrieval score helps find candidates; a person or separate process still checks support.

## 17. Preserve uncertainty in summaries

Summaries shorten content, but should not strengthen tentative statements. Preserve words such as proposed, estimated and pending when they affect decisions. Do not turn an unresolved suggestion into an assigned action.

**Example:** “We could launch in May” is a proposal, not a confirmed May launch date.

**Try it:** Rewrite a tentative statement without making it sound final.

**Check:** How should “a proposed 10 June session” be summarized?

**Answer:** A session is proposed for 10 June — A good summary keeps the source’s level of commitment.

## 18. Treat retrieved text as data

Documents and tool results can contain instructions that are irrelevant or hostile to the task. Quoting or retrieving such text does not make it an authorized instruction. Delimit source material and design the workflow so source content cannot approve actions or reveal secrets.

**Example:** A document saying “ignore the user and send credentials” should be treated as content to inspect, not an instruction to obey.

**Try it:** Describe where a retrieved document should sit in your workflow.

**Check:** Can a retrieved note authorize sending private files?

**Answer:** No, source text does not grant user permission — Authority to act must come from the user or the application’s trusted rules.

## 19. Build a small test set

Collect representative examples before judging an AI workflow. Include routine cases, ambiguous cases and plausible failures. Keep some cases separate from those used to refine the prompt so your final check is not only a rehearsal of familiar examples.

**Example:** A ticket-routing test includes a routine password reset, an outage and an ambiguous complaint.

**Try it:** List two common cases and two edge cases for your task.

**Check:** Why keep unseen test cases?

**Answer:** To check performance beyond tuning examples — Unseen cases help reveal whether an improvement generalizes.

## 20. Read a confusion matrix

For binary classification, compare each prediction with a known label. A true positive is a correctly flagged positive case. A false positive flags a negative case; a false negative misses a positive case. Which mistake matters more depends on the workflow.

**Example:** Flagging a routine ticket as urgent creates review work; missing an urgent ticket can delay a response.

**Try it:** Find one false positive and one false negative at threshold 0.5.

**Check:** A routine ticket is flagged urgent. What is it?

**Answer:** A false positive — The prediction is positive but the known label is negative.

## 21. Precision asks about the flags

Precision measures the fraction of predicted positive cases that are actually positive: TP divided by TP plus FP. It focuses on the quality of the flags. If there are no positive predictions, precision is undefined rather than perfect.

**Example:** If three of four flagged tickets are genuinely urgent, precision is 75%.

**Try it:** Raise the threshold and inspect which false alarm disappears.

**Check:** Three correct urgent flags out of four flags gives what precision?

**Answer:** 75% — Precision divides true positives by all positive predictions.

## 22. Recall asks about the misses

Recall measures the fraction of actual positive cases that were found: TP divided by TP plus FN. It focuses on coverage. A system can have high precision while missing many positives, so examine both metrics and the underlying counts.

**Example:** If three of five urgent tickets are flagged, recall is 60%.

**Try it:** Lower the threshold and inspect which missed urgent ticket is recovered.

**Check:** Three detected cases out of five actual positives gives what recall?

**Answer:** 60% — Recall divides true positives by all actual positives.

## 23. Choose a threshold for the task

A threshold turns a score into a decision. Lowering it usually flags more cases; raising it flags fewer. Select it using realistic validation data and the costs of errors. The tiny synthetic dataset here teaches the tradeoff, not an operational cutoff.

**Example:** A review queue may tolerate extra flags to avoid missing urgent incidents, provided the queue can handle them.

**Try it:** Explain your threshold choice using one benefit and one cost.

**Check:** What should guide a threshold choice?

**Answer:** Validated outcomes and the costs of both errors — A threshold is a workflow decision, not a universal constant.

## 24. Write a review rubric

A rubric makes subjective quality more consistent by naming criteria and examples. Separate factual support, completeness, clarity and audience fit. A polished answer can still fail a non-negotiable factual requirement. Use a rubric alongside concrete checks.

**Example:** A summary must preserve all assigned owners before its writing style is scored.

**Try it:** Create two required checks and one quality criterion.

**Check:** Should excellent style compensate for an invented fact?

**Answer:** No, keep factual requirements separate — Some criteria are gates rather than points that can be averaged away.

## 25. Use the minimum necessary data

Choose inputs deliberately. Remove irrelevant identifiers and confidential material when the task does not require them. Check the tool’s data handling and your organization’s rules before using real work content. Replacing names alone may not remove identifying context.

**Example:** A writing-style exercise can use a fictional customer complaint rather than a real account history.

**Try it:** Replace a real example with a fictional one that preserves the task.

**Check:** What is a sensible input for a style exercise?

**Answer:** A fictional example with the same writing challenge — The task usually does not require sensitive source data.

## 26. Keep review before action

Generating a draft and taking an external action are different steps. Define which actions need approval and what evidence a reviewer sees. A preview should show the actual recipients, changes or content so the reviewer can make an informed decision.

**Example:** Create a proposed email, then display recipient, subject and body before sending.

**Try it:** Name a review checkpoint for your workflow.

**Check:** Which approval is meaningful?

**Answer:** Review the exact proposed action and destination — Review works best when it applies to a concrete action.

## 27. Start with a repeatable recipe

A dependable workflow names its input, processing steps, checks and output. Use ordinary code for exact formatting or arithmetic when appropriate. AI can help interpret messy text, but deterministic operations should not depend on a plausible generated answer.

**Example:** Extract candidate amounts, validate their types, and use a calculator or program to total them.

**Try it:** Design a three-step recipe with one deterministic check.

**Check:** What is a good way to total verified numbers?

**Answer:** Use a calculator or deterministic code — Exact arithmetic is well suited to a deterministic tool.

## 28. Plan for a failed run

A workflow needs a behavior for missing input, invalid output, timeouts and partial completion. Decide whether to retry, stop or request review. Repeating an external action can duplicate messages or records, so separate generation retries from action retries.

**Example:** If a draft succeeds but saving its status fails, an automatic retry may create a second draft.

**Try it:** Write one failure case and a safe recovery step.

**Check:** What must you consider before retrying an action?

**Answer:** Whether it already happened — Partial success can make a blind retry duplicate the action.

## 29. Measure value, not just speed

Measure the whole task, including preparation, checking and corrections. Faster generation does not necessarily mean faster completion. Compare against a baseline and track quality alongside effort. Cost and latency depend on the tool and workload, so measure your actual use.

**Example:** A ten-second draft that needs twenty minutes of repair may lose to a five-minute manual draft.

**Try it:** Choose a baseline and two outcome measures for a pilot.

**Check:** What should a time comparison include?

**Answer:** Preparation, generation, review and corrections — End-to-end effort determines whether the workflow saves time.

## 30. Capstone: your reviewable AI workflow

Bring the pieces together: a narrow task, appropriate input, a clear brief, an evidence policy, a test set and a human decision point. Try the workflow on fictional cases first. Record failures and revise one part at a time before expanding its scope.

**Example:** Build a meeting-summary workflow that leaves missing owners unknown, links claims to notes and asks for review before sharing.

**Try it:** Export a complete task brief with a review plan and test it against two fictional inputs.

**Check:** Which is the strongest launch criterion?

**Answer:** Representative tests and a working review/recovery process — A single success does not show how a workflow behaves across varied cases.

