No take-home mathematics question is permanently “AI-proof”. A question that one model mishandles today may be solved after a model update, a different prompt, access to tools or a student follow-up. Designing around a secret trick therefore creates false confidence.
A more defensible goal is AI-resistant assessment: tasks that validly assess the intended HSC mathematics, make a student's reasoning visible and gather enough evidence over time for a teacher to judge authenticity. This is an assessment-design problem, not a detector problem.
NESA's current position, checked 3 August 2026, is that unapproved AI use in assignments breaches academic integrity, work must be the student's own or appropriately acknowledged, and schools provide specific advice about permitted use. NESA also advises varied assessment at regular intervals and says teachers are best placed to recognise student voice and judge authenticity. Read the official NESA advice on student use of AI alongside your sector and school policy.
Download the AI-resistant maths assessment checklist (PDF). Use it during task drafting to record the intended construct, permitted AI conditions, process evidence, verification method, adjustments and benchmark status.
Begin With the Construct
Before trying to make a task resistant to outsourcing, state what the task is meant to measure.
For example:
- selecting a model from information in context
- carrying out a mathematical process accurately
- interpreting a result and its limitations
- constructing a proof or justification
- comparing methods
- using technology appropriately
Then ask which evidence must come from the student. If method selection is central, a correct final number is weak evidence. If interpretation is central, pages of algebra may be beside the point.
AI resistance must not damage validity. An obscure context, deliberate ambiguity or excessive reading load may make a question harder for everyone while measuring less mathematics.
Use an Evidence Bundle
A single polished submission is fragile evidence. A stronger task collects several connected pieces:
- Supervised start: students interpret the problem, define variables or propose a plan in class.
- Working record: students retain calculations, revisions, data decisions and unsuccessful approaches.
- Final response: students communicate the resolved mathematics.
- Verification: a short conference, changed-value question or in-class reflection checks ownership.
- Earlier comparison: the teacher compares the work with routine classroom evidence.
The aim is not surveillance. It is to make mathematical development observable and to give students legitimate opportunities to explain their decisions.
Question Patterns That Strengthen Authenticity Evidence
Ask for a decision before a calculation
Require students to choose a model, identify assumptions or reject irrelevant information. Mark the decision separately from execution.
Use local, teacher-controlled data carefully
A small dataset collected in class can create a shared evidence trail. Remove personal information, check permissions and provide an equivalent dataset for absent students. Novel data alone does not prevent AI use; its value is that the class has discussed its collection and limitations.
Add a perturbation
After submission, change one value or constraint and ask what changes without repeating the whole solution. A student who owns the model should be able to trace the effect.
Ask for error diagnosis
Provide a plausible solution with one consequential error. Ask students to identify the first invalid step, explain its effect and repair it. Ensure the error targets the intended mathematics rather than a trick in wording.
Compare representations or methods
Ask when a graph, recurrence, table, formula or simulation is most useful. The explanation can reveal understanding that a final answer hides.
Include a brief oral check
Use two consistent prompts: “Why did you choose this method?” and “What would change if this condition changed?” Keep a concise record and apply the same process fairly under school policy.
Three Concrete Before-and-After Rewrites
The goal of a rewrite is to collect better evidence of the same mathematics. It is not to make the language strange or to add an unrelated presentation hurdle.
Rewrite 1: compound interest model selection
Before: “$18,000 is invested at 5.4% per annum, compounded monthly, for four years. Calculate the final value.”
This is useful fluency practice, but a polished final response gives limited evidence about who selected the model or understood the compounding period.
After: “Two students propose models for the investment: Model A applies 5.4% once per year; Model B applies one-twelfth of the annual rate each month. Before calculating, select the model that matches the stated conditions and explain the period used. Calculate the final value. In a two-minute supervised follow-up, explain what changes if compounding becomes quarterly while the nominal annual rate stays the same.”
The after version separately records model selection, calculation and transfer. It does not become “AI-proof”; a tool can still solve it. The teacher gains connected evidence from a notified supervised checkpoint and a later response.
Rewrite 2: statistical comparison
Before: “Calculate the mean and standard deviation of each dataset and identify which is more consistent.”
That prompt may reward routine calculation while hiding whether the student understands consistency or the influence of an unusual value.
After: “Using the two class datasets issued in the supervised lesson, predict which group is more consistent and identify the feature that informed your prediction. Calculate suitable statistics, compare them with the prediction and justify a conclusion. Your final submission must include the annotated plot produced in class. At verification, one value will be changed; explain whether the conclusion is robust without recomputing every statistic.”
The local dataset is not a security trick. Its value is the contemporaneous discussion and retained plot. Provide an equivalent teacher-controlled dataset for an absent student, avoid personal data and keep the verification proportional.
Rewrite 3: optimisation and error diagnosis
Before: “Find the dimensions of the rectangle with maximum area under the curve.”
The task may validly assess calculus, but a single neat solution says little about how the student checked the domain, stationary point or conclusion.
After: “A proposed solution differentiates the area function correctly but accepts a stationary point outside the physical domain. Identify the first step at which the conclusion becomes invalid, repair the argument and state the feasible dimensions. Then sketch or describe a second check that supports the maximum. During the supervised start, record the domain before using calculus.”
This keeps the construct mathematical: domain, differentiation, maximum and justification. The added evidence is a before-calculation decision and a response to a plausible error, not a demand for decorative prose.
For each rewrite, update the marking guideline at the same time. If selection, explanation or verification matters, assign it an explicit criterion. Otherwise the added component becomes workload without assessment value.
Draw a Clear Boundary: Homework Practice vs Formal Assessment
Routine questions still belong in teaching. A student benefits from repeated substitution, algebraic fluency and immediate feedback. If the purpose is formative homework practice, a faculty may permit named AI uses—such as checking a completed answer, generating a similar example or explaining an error—under school policy. The teacher can then ask students to retain their first attempt, identify what help was used and correct the mathematics.
Formal assessment is different because the school uses the evidence to judge achievement. The notification must state whether AI is prohibited, permitted for limited purposes or required; what must be acknowledged; which work must occur under supervision; and how authenticity may be checked. Do not rely on a general faculty slogan if the task-specific conditions are ambiguous.
The boundary should follow purpose, not location. A take-home task can be formal assessment, and an in-class quiz can be formative. Label resources and learning-management-system posts clearly so students are not asked to infer the rule from where the file appears.
If AI is permitted during practice but prohibited for a formal response, teach the transition. Show students how to close the tool, begin from a fresh task, retain required working and seek clarification. A coherent policy does not punish students for using a practice method that the teacher appeared to encourage.
Faculty Action: Test One Question, Not Your Whole Bank
Choose a high-weight question from your next draft. Adapt a starting point from HSC questions by topic, identify the construct, add one process checkpoint and write two verification prompts. Then use the benchmark protocol below before approving the task.
Benchmark Results
No reproducible live-model benchmark run has been completed or published for this article as at 3 August 2026. Consequently, this article withholds claims about model accuracy, failure rates and whether any example is resistant to a named product. The protocol and table below are provided so a future run can be published transparently instead of reconstructed after seeing the outputs.
| Field | Result to publish | |---|---| | Run date and time | Not run as at 3 August 2026 | | Question-set version and checksum | To be recorded before testing | | Model, product surface and displayed version | To be recorded for every condition | | Tools available | Browsing, code, calculator, file or image access stated explicitly | | Prompt conditions and repetitions | Exact prompts and run count linked | | Blind marking method | Criteria, marker count and disagreement process stated | | Marks by criterion | Withheld until a reproducible run is completed | | Mathematical and factual error categories | Withheld until a reproducible run is completed | | Follow-up verification performance | Withheld until a reproducible run is completed | | Limitations | Date, sample size, tool settings and generalisation limits |
Do not replace “not run” cells with an informal classroom impression. If the question set, prompts or raw outputs cannot be shared because they remain live assessment material, publish the protocol and aggregate results only after the task, explain the restriction and retain the evidence securely for internal review.
Reproducible live-model protocol
The protocol below is a faculty method for testing its own questions. Results would be a dated snapshot, not a permanent security rating.
1. Freeze the test materials
Save:
- exact question wording and all diagrams or data
- marking guideline and intended construct
- model name and version shown by the vendor
- product or API surface used
- tool access enabled, including browsing, code or calculator tools
- date, locale and relevant settings
Model names alone are insufficient because product tools and configurations can change performance.
2. Pre-register prompts
Use at least three prompt conditions:
- the question alone
- a neutral instruction to show full reasoning
- a realistic student-style request for help
Do not add corrective hints after seeing a weak first answer unless “with hints” is a separate condition. Save every prompt exactly.
3. Repeat runs
Generative outputs vary. Run each condition several times in fresh sessions and record all outputs, including refusals or tool errors. A single failure is not evidence that the question is resistant.
4. Blind-mark against the same criteria
Remove model identifiers and have at least two teachers independently apply the task's marking guideline. Record:
- marks by criterion
- mathematical errors
- whether the model disclosed assumptions
- whether citations, data or calculations were fabricated
- disagreements between markers
Resolve marker differences using the same moderation process intended for student work.
5. Test the verification prompt
Give the model the same changed-value or oral-style follow-up planned for students. Record whether it can maintain a coherent solution. A weak follow-up may add inconvenience without adding useful evidence.
6. Interpret cautiously
Report results only for the tested date, versions, settings and prompts. Do not generalise from one commercial chatbot to “AI”. Do not publish the test set before the live task. Re-test after material model or tool updates if the result informs a high-stakes design decision.
If your school permits vendor testing, do not enter student work, names or other personal or sensitive information into public AI tools. NESA's June 2026 AI transparency statement concerns NESA's own operations rather than school requirements, but its explicit separation of public tools and personal information is a useful risk reminder. Follow your own approved privacy and procurement rules.
What Not to Claim
Avoid claims such as:
- “This question cannot be solved by ChatGPT.”
- “A detector proves the response was AI-written.”
- “Unusual wording makes the task secure.”
- “A handwritten response must be authentic.”
- “Every student must disclose any AI use” unless that is actually the notified school requirement for this task.
Instead, say what the design does: “The task collects supervised planning, a working record and a short verification response.” Those are observable features.
Notify Students Precisely
State, in plain language:
- whether AI is prohibited, permitted for named purposes or required
- which stages of the task the rule covers
- what acknowledgement is required for permitted use
- what process evidence must be retained
- how authenticity may be checked
- where students can ask before using a tool
- which school malpractice and appeal procedures apply
Do not introduce a material restriction after students start. Align the notification with current NESA rules, sector requirements and the school's assessment policy. The 2026 HSC guide names unauthorised generative-AI use within its academic-integrity advice; see Before you start your HSC.
Preserve Access and Fairness
An oral verification step can disadvantage students if it becomes an unnotified second assessment or ignores approved adjustments. Keep it short, linked to the submitted mathematics and governed by a common protocol. Plan access arrangements through the school's existing process.
Likewise, requiring extensive handwritten drafts may measure speed, transcription or executive function more than the target outcome. Collect only the process evidence needed for a sound judgement.
A Faculty Review Checklist
Before approval, ask:
- Is the intended mathematical construct explicit?
- Does each task component gather relevant evidence?
- Are the AI conditions clear before commencement?
- Can permitted and prohibited assistance be distinguished?
- Is there a supervised or contemporaneous evidence point?
- Are verification prompts consistent and proportionate?
- Have privacy and access requirements been addressed?
- Has the marking guideline been tested on alternative valid methods?
- If an AI benchmark was run, are the prompts, versions, tools, repeats and limits recorded?
The best defence is not a one-off “AI-hard” question. It is a coherent assessment program that sees students solve, explain, revise and transfer mathematics across time. Pair this guide with writing an HSC maths assessment task and structuring a marking rubric.