Toolkit / Assignment Package: AI-Permitted Work + Complementary Evidence
Assignment Package: AI-Permitted Work + Complementary Evidence
A proportionate assessment-twins pattern combining disclosed AI-assisted work with a short, independently completed application or explanation.
Use this pattern when AI-assisted performance is itself relevant but the final artifact cannot, on its own, support the judgment you need to make about a student’s understanding. Pair the disclosed work product with a short, independently completed task aimed at the same learning outcome through different evidence.
The complement may be a brief explanation, a new fact variation, an in-class application, a source defense, or a structured conference. It need not be oral, and it should not become a second full assessment.
Purpose
The design separates two warranted questions:
- Can the student produce and supervise high-quality AI-assisted legal work?
- Can the student personally explain, apply, or defend the underlying legal judgment?
Use the pairing only when both inferences matter. If the work product already supplies sufficient evidence, do not add a complement merely as surveillance.
Appropriate courses
- Seminars, capstones, drafting, negotiation, and policy projects permitting broad AI assistance
- Doctrinal take-home problems where professional tool use and individual application both matter
- Simulations assessing a product plus advocacy, counseling, or decision explanation
- Clinics and externships only with specific supervisory authorization and protected information
Learning outcomes
Students should be able to:
- produce an accurate, well-supported work product while supervising permitted assistance;
- verify authority, facts, and consequential claims;
- explain the product’s central legal and professional choices;
- apply the same governing framework to a modest variation;
- identify uncertainty and limits; and
- disclose assistance accurately.
Assignment sequence
- Announce both components: publish the work-product and complementary-task criteria together.
- Complete AI-permitted work: students may use approved AI assistance within stated data and disclosure limits.
- Submit: collect the product and concise use statement.
- Complete the complement: use a short no-AI explanation or application under equivalent conditions.
- Evaluate jointly: interpret each component for the outcome it actually measures; define in advance how inconsistent evidence will be handled.
- Follow up proportionately: use ordinary feedback or a targeted clarification, not an automatic misconduct inference.
Choose the least burdensome complement
| Complement | Best use | Caution |
|---|---|---|
| 300-word explanation | Central choices, authority, and verification | Can become formulaic if prompts are generic |
| New fact variation | Transfer of doctrine or strategy | Keep difficulty comparable across students |
| Source defense | Research judgment and authority use | Do not reward database access or memorized citation detail |
| Five-minute conference | Reasoning, audience, professional judgment | Requires calibration, scheduling, practice, and alternatives |
| In-class application | Scalable individual evidence | Ensure accommodations and avoid turning it into an unrelated speed test |
| Implementation sample | Review workload, scoring, and question quality | Not evidence of unsampled students’ competence and not a substitute for an individually graded complement |
Student-facing instructions
You may use generative AI for the work product within the limits below. You remain responsible for the facts, authorities, analysis, confidentiality, and professional judgment in the submission, and you must disclose consequential assistance.
You will also complete a short individual task addressing the same learning outcome. The task may ask you to explain a central decision, apply the governing framework to a new fact, or defend a source or recommendation. It is not a memory test about your exact wording. Its purpose is to provide complementary evidence of your understanding.
The individual task must be completed without AI unless the instructions expressly say otherwise. The criteria and relationship between the two components appear in the rubric.
AI rule
For the work product, list permitted activities and covered stages. State whether students may upload the prompt, course materials, sources, or drafts, and which tools are institutionally appropriate. For the complement, specify the independent conditions and permitted assistive technology.
Never require students to enter their work into an unlicensed AI tool. A course rule does not override exam instructions, supervision, clinic or placement rules, client duties, or current Penn data guidance.
Disclosure
I used [tool] for [brainstorming/research leads/outlining/drafting/critique/revision/other]. The assistance materially affected [parts or decisions]. I independently verified authorities, quotations, facts, and consequential claims by [method]. I remain responsible for the submitted product.
Do not make transcript length a quality metric. Ask for excerpts only when needed for a stated AI-supervision outcome.
Deliverables
- AI-assisted work product
- Concise AI-use and verification statement
- Announced complementary task
- Optional targeted clarification if the evidence materially conflicts
Evaluation rubric
| Dimension | Weight | Successful evidence |
|---|---|---|
| Work-product legal quality | 35% | Accurate authority and facts, sound analysis, appropriate judgment, and effective communication. |
| Supervision and verification | 20% | Assistance is used within bounds; consequential content is checked; limitations and uncertainty are recognized. |
| Individual explanation or application | 30% | Student accurately explains or transfers the central framework and can justify important choices. |
| Integration of judgment | 10% | Product and complement reflect a coherent understanding; student can identify what would change the conclusion. |
| Disclosure and compliance | 5% | Use statement is specific, candid, and complete. |
Set a transparent consequence for weak or inconsistent evidence. Options include criterion-specific feedback, a limited follow-up, revision, or separate component grades. Do not treat inconsistency alone as proof of misconduct.
Class-size variants
Small class: five- to eight-minute structured conferences using two common questions and one individualized follow-up.
Medium class: written explanation plus a fact variation completed in class; conferences only where evidence is ambiguous.
Large class: give every student a short common written application when individual competence is graded. Use teaching-assistant scoring only with scripts, calibration examples, and faculty review. A sample of oral follow-ups may inform implementation review or evaluator calibration, but it is not evidence of unsampled students’ competence.
Accessibility and equity
- Oral performance is not the default. Choose written, asynchronous, or performance alternatives that assess the same outcome.
- If oral work is necessary, give practice, questions or domains in advance where appropriate, transparent criteria, sufficient response time, and approved accommodations.
- Consider disability, anxiety, language background, time zones, technology, and cultural differences in live interaction.
- Train and calibrate multiple evaluators; record scores and reasons consistently.
- Keep the complement proportionate. Added workload and stress can undermine the validity the pairing is meant to improve.
Evidence and limitations
Roe, Perkins & Giray’s “Assessment Twins” offers the principal validity framework, but the authors expressly state that it has not yet been empirically validated and caution that twinning is not appropriate in every context. Cong-Lem’s teacher-led interview study reports reflections from 24 master’s-level EFL students, not comparative validity or authorship-detection outcomes.
A systematic review of oral-assessment performance identified structured practice, timely feedback, technology supports, and self-reflection as facilitators. It does not establish oral assessment as universally fair or AI-proof. See Stephenson, Johnson-Glauch & Cruchley, “Interventions and Facilitators of Oral Assessment Performance in Higher Education”.
This package is therefore a promising design with limited direct validation. Pilot it at low stakes, review student experience and scoring consistency, and retain it only when the complementary evidence improves the decision.