Toolkit / Assignment Package: Audit and Correct AI Legal Work
Assignment Package: Audit and Correct AI Legal Work
A complete exercise for assessing legal analysis, authority verification, and professional judgment by having students diagnose and repair AI output.
Use this pattern when the learning goal is to verify, critique, and improve assisted legal work. Give every student the same AI-generated response or a controlled set of variants. Students audit the output against the record and governing authority, then submit corrections and a defensible replacement.
The object being graded is the student’s diagnosis and judgment, not the apparent quality of the machine response.
Purpose
This assignment turns predictable AI failure modes into an assessment of legal capability. It can reveal whether students can distinguish fluent language from supported analysis, locate controlling authority, identify missing facts and counterarguments, and make a professional recommendation.
Appropriate courses
- Doctrinal courses after students have learned the governing framework
- Legal research and writing, advocacy, transactional drafting, and negotiation
- Professional responsibility and clinics using fictional or properly authorized material
- Introductory AI-literacy sessions tied to a substantive legal problem
For high-stakes use, give students enough time and source access to evaluate the legal claims rather than merely hunt for invented citations.
Learning outcomes
Students should be able to:
- test an answer against the relevant facts, issues, law, and requested task;
- locate and read every authority on which a consequential claim depends;
- distinguish false citation, inaccurate characterization, incomplete analysis, and weak judgment;
- explain why a defect matters to the reader or client;
- correct the response without importing new unsupported claims; and
- state the limits of the revised answer.
Assignment sequence
- Orient: provide the problem, source set or research boundary, AI output, and evaluation criteria.
- Inventory claims: students identify the output’s legal, factual, and professional assertions.
- Verify: students locate authorities and compare propositions, quotations, dates, posture, and treatment.
- Diagnose: students annotate defects and omissions by category and consequence.
- Correct: students prepare a revised answer or correction memorandum.
- Explain: students identify the highest-risk defect and the verification step that changed their analysis.
Student-facing instructions
The attached response was produced with generative AI. Do not assume that any statement, citation, quotation, or conclusion is correct. Audit the response as if you were responsible for the work before it reached a supervising lawyer, court, client, or other decision-maker.
For each material defect, identify (1) the claim, (2) the type of defect, (3) the source used to check it, (4) why the defect matters, and (5) the required correction. Then submit a corrected response that answers the assigned question. Your score depends on your legal analysis, verification, prioritization, and correction—not on how many small errors you list.
AI rule
Choose one of two designs and state it on the assignment:
- Controlled audit: students receive instructor-provided AI output and may not ask another AI system to perform the audit. This best isolates their verification and analysis.
- Supervised comparison: after completing an initial audit, students may use an approved tool to seek a second critique, then identify what it caught, missed, or misstated.
In either design, prohibit sensitive or restricted information and require direct source verification. AI output is never legal authority.
Disclosure
For a controlled audit, a simple compliance statement is enough. For a supervised comparison:
I used [tool] after completing my initial audit to [purpose]. It identified [brief description]. I independently verified or rejected those suggestions using [sources/method]. AI-generated wording [does/does not] appear in my corrected response.
Deliverables
- Annotated AI output or defect table
- Source-verification record with stable citations or links
- Corrected response
- Short risk-and-judgment note (suggested length: 250–400 words)
- AI-use statement if students use a second tool
Evaluation rubric
| Dimension | Weight | Successful evidence |
|---|---|---|
| Issue and defect diagnosis | 25% | Identifies material errors and omissions across law, facts, reasoning, task fit, and professional context; prioritizes consequential defects. |
| Authority verification | 25% | Locates reliable sources; checks existence, proposition, quotation, posture, date, and current treatment as appropriate. |
| Corrected legal analysis | 30% | Replacement accurately applies governing authority to the facts, addresses counterarguments, and answers the assigned question. |
| Professional judgment | 15% | Explains consequences, uncertainty, reader/client needs, and what requires escalation or further research. |
| Communication and compliance | 5% | Audit trail is usable, concise, and complete; required disclosure is accurate. |
Do not award points for finding trivia. A shorter audit that catches the conclusion-changing defect can be stronger than a long list of stylistic complaints.
Class-size variants
Small class: use different outputs for teams; hold a debrief in which students defend correction priorities.
Medium class: assign a common output; divide verification categories among groups before individual corrected submissions.
Large class: use a bounded source packet and structured defect table; score a small set of high-value criteria; use calibrated examples for teaching assistants. An in-class polling debrief can surface disagreements without requiring individual oral exams.
Accessibility and equity
- Make the source set available in accessible formats and allow ordinary accommodations and assistive research tools.
- Avoid grading speed unless rapid verification is an explicit, justified outcome.
- Teach the verification method before assessing it; prior familiarity with premium legal AI tools should not be an unstated advantage.
- Use fictional, public, or authorized facts. Do not upload client, student, or nonpublic materials to generate the exercise.
- If deliberately flawed output includes biased or harmful content, warn students and explain its pedagogical necessity.
Evidence and limitations
Research repeatedly shows that language models can produce plausible but unreliable legal output, and course grounding does not eliminate error. Those studies support the need for verification; they do not directly prove that this particular exercise improves student learning. Treat the package as an aligned design recommendation and evaluate it in your course.
Useful legal sources include Choi, “Off-the-Shelf Large Language Models Are Unreliable Judges”, and Ouellette et al., “Can AI Hold Office Hours?”. The latter tested 185 questions in one patent-law corpus using October 2024 systems, so its error rates should not be generalized to every current tool or subject.