HomeBlogAI-Resistant HSC Maths Assessment Questions
Published 3 August 2026·By Emmanuel Koutsouklakis·11 min read

AI-Resistant HSC Maths Assessment Questions

Design HSC maths assessment that produces stronger authenticity evidence, plus a reproducible protocol for testing questions against AI tools.

No take-home mathematics question is permanently “AI-proof”. A question that one model mishandles today may be solved after a model update, a different prompt, access to tools or a student follow-up. Designing around a secret trick therefore creates false confidence.

A more defensible goal is AI-resistant assessment: tasks that validly assess the intended HSC mathematics, make a student's reasoning visible and gather enough evidence over time for a teacher to judge authenticity. This is an assessment-design problem, not a detector problem.

NESA's current position, checked 3 August 2026, is that unapproved AI use in assignments breaches academic integrity, work must be the student's own or appropriately acknowledged, and schools provide specific advice about permitted use. NESA also advises varied assessment at regular intervals and says teachers are best placed to recognise student voice and judge authenticity. Read the official NESA advice on student use of AI alongside your sector and school policy.

Download the AI-resistant maths assessment checklist (PDF). Use it during task drafting to record the intended construct, permitted AI conditions, process evidence, verification method, adjustments and benchmark status.


Begin With the Construct

Before trying to make a task resistant to outsourcing, state what the task is meant to measure.

For example:

Then ask which evidence must come from the student. If method selection is central, a correct final number is weak evidence. If interpretation is central, pages of algebra may be beside the point.

AI resistance must not damage validity. An obscure context, deliberate ambiguity or excessive reading load may make a question harder for everyone while measuring less mathematics.

Use an Evidence Bundle

A single polished submission is fragile evidence. A stronger task collects several connected pieces:

  1. Supervised start: students interpret the problem, define variables or propose a plan in class.
  2. Working record: students retain calculations, revisions, data decisions and unsuccessful approaches.
  3. Final response: students communicate the resolved mathematics.
  4. Verification: a short conference, changed-value question or in-class reflection checks ownership.
  5. Earlier comparison: the teacher compares the work with routine classroom evidence.

The aim is not surveillance. It is to make mathematical development observable and to give students legitimate opportunities to explain their decisions.

Question Patterns That Strengthen Authenticity Evidence

Ask for a decision before a calculation

Require students to choose a model, identify assumptions or reject irrelevant information. Mark the decision separately from execution.

Use local, teacher-controlled data carefully

A small dataset collected in class can create a shared evidence trail. Remove personal information, check permissions and provide an equivalent dataset for absent students. Novel data alone does not prevent AI use; its value is that the class has discussed its collection and limitations.

Add a perturbation

After submission, change one value or constraint and ask what changes without repeating the whole solution. A student who owns the model should be able to trace the effect.

Ask for error diagnosis

Provide a plausible solution with one consequential error. Ask students to identify the first invalid step, explain its effect and repair it. Ensure the error targets the intended mathematics rather than a trick in wording.

Compare representations or methods

Ask when a graph, recurrence, table, formula or simulation is most useful. The explanation can reveal understanding that a final answer hides.

Include a brief oral check

Use two consistent prompts: “Why did you choose this method?” and “What would change if this condition changed?” Keep a concise record and apply the same process fairly under school policy.

Three Concrete Before-and-After Rewrites

The goal of a rewrite is to collect better evidence of the same mathematics. It is not to make the language strange or to add an unrelated presentation hurdle.

Rewrite 1: compound interest model selection

Before: “$18,000 is invested at 5.4% per annum, compounded monthly, for four years. Calculate the final value.”

This is useful fluency practice, but a polished final response gives limited evidence about who selected the model or understood the compounding period.

After: “Two students propose models for the investment: Model A applies 5.4% once per year; Model B applies one-twelfth of the annual rate each month. Before calculating, select the model that matches the stated conditions and explain the period used. Calculate the final value. In a two-minute supervised follow-up, explain what changes if compounding becomes quarterly while the nominal annual rate stays the same.”

The after version separately records model selection, calculation and transfer. It does not become “AI-proof”; a tool can still solve it. The teacher gains connected evidence from a notified supervised checkpoint and a later response.

Rewrite 2: statistical comparison

Before: “Calculate the mean and standard deviation of each dataset and identify which is more consistent.”

That prompt may reward routine calculation while hiding whether the student understands consistency or the influence of an unusual value.

After: “Using the two class datasets issued in the supervised lesson, predict which group is more consistent and identify the feature that informed your prediction. Calculate suitable statistics, compare them with the prediction and justify a conclusion. Your final submission must include the annotated plot produced in class. At verification, one value will be changed; explain whether the conclusion is robust without recomputing every statistic.”

The local dataset is not a security trick. Its value is the contemporaneous discussion and retained plot. Provide an equivalent teacher-controlled dataset for an absent student, avoid personal data and keep the verification proportional.

Rewrite 3: optimisation and error diagnosis

Before: “Find the dimensions of the rectangle with maximum area under the curve.”

The task may validly assess calculus, but a single neat solution says little about how the student checked the domain, stationary point or conclusion.

After: “A proposed solution differentiates the area function correctly but accepts a stationary point outside the physical domain. Identify the first step at which the conclusion becomes invalid, repair the argument and state the feasible dimensions. Then sketch or describe a second check that supports the maximum. During the supervised start, record the domain before using calculus.”

This keeps the construct mathematical: domain, differentiation, maximum and justification. The added evidence is a before-calculation decision and a response to a plausible error, not a demand for decorative prose.

For each rewrite, update the marking guideline at the same time. If selection, explanation or verification matters, assign it an explicit criterion. Otherwise the added component becomes workload without assessment value.

Draw a Clear Boundary: Homework Practice vs Formal Assessment

Routine questions still belong in teaching. A student benefits from repeated substitution, algebraic fluency and immediate feedback. If the purpose is formative homework practice, a faculty may permit named AI uses—such as checking a completed answer, generating a similar example or explaining an error—under school policy. The teacher can then ask students to retain their first attempt, identify what help was used and correct the mathematics.

Formal assessment is different because the school uses the evidence to judge achievement. The notification must state whether AI is prohibited, permitted for limited purposes or required; what must be acknowledged; which work must occur under supervision; and how authenticity may be checked. Do not rely on a general faculty slogan if the task-specific conditions are ambiguous.

The boundary should follow purpose, not location. A take-home task can be formal assessment, and an in-class quiz can be formative. Label resources and learning-management-system posts clearly so students are not asked to infer the rule from where the file appears.

If AI is permitted during practice but prohibited for a formal response, teach the transition. Show students how to close the tool, begin from a fresh task, retain required working and seek clarification. A coherent policy does not punish students for using a practice method that the teacher appeared to encourage.


Faculty Action: Test One Question, Not Your Whole Bank

Choose a high-weight question from your next draft. Adapt a starting point from HSC questions by topic, identify the construct, add one process checkpoint and write two verification prompts. Then use the benchmark protocol below before approving the task.


Benchmark Results

No reproducible live-model benchmark run has been completed or published for this article as at 3 August 2026. Consequently, this article withholds claims about model accuracy, failure rates and whether any example is resistant to a named product. The protocol and table below are provided so a future run can be published transparently instead of reconstructed after seeing the outputs.

| Field | Result to publish | |---|---| | Run date and time | Not run as at 3 August 2026 | | Question-set version and checksum | To be recorded before testing | | Model, product surface and displayed version | To be recorded for every condition | | Tools available | Browsing, code, calculator, file or image access stated explicitly | | Prompt conditions and repetitions | Exact prompts and run count linked | | Blind marking method | Criteria, marker count and disagreement process stated | | Marks by criterion | Withheld until a reproducible run is completed | | Mathematical and factual error categories | Withheld until a reproducible run is completed | | Follow-up verification performance | Withheld until a reproducible run is completed | | Limitations | Date, sample size, tool settings and generalisation limits |

Do not replace “not run” cells with an informal classroom impression. If the question set, prompts or raw outputs cannot be shared because they remain live assessment material, publish the protocol and aggregate results only after the task, explain the restriction and retain the evidence securely for internal review.

Reproducible live-model protocol

The protocol below is a faculty method for testing its own questions. Results would be a dated snapshot, not a permanent security rating.

1. Freeze the test materials

Save:

Model names alone are insufficient because product tools and configurations can change performance.

2. Pre-register prompts

Use at least three prompt conditions:

Do not add corrective hints after seeing a weak first answer unless “with hints” is a separate condition. Save every prompt exactly.

3. Repeat runs

Generative outputs vary. Run each condition several times in fresh sessions and record all outputs, including refusals or tool errors. A single failure is not evidence that the question is resistant.

4. Blind-mark against the same criteria

Remove model identifiers and have at least two teachers independently apply the task's marking guideline. Record:

Resolve marker differences using the same moderation process intended for student work.

5. Test the verification prompt

Give the model the same changed-value or oral-style follow-up planned for students. Record whether it can maintain a coherent solution. A weak follow-up may add inconvenience without adding useful evidence.

6. Interpret cautiously

Report results only for the tested date, versions, settings and prompts. Do not generalise from one commercial chatbot to “AI”. Do not publish the test set before the live task. Re-test after material model or tool updates if the result informs a high-stakes design decision.

If your school permits vendor testing, do not enter student work, names or other personal or sensitive information into public AI tools. NESA's June 2026 AI transparency statement concerns NESA's own operations rather than school requirements, but its explicit separation of public tools and personal information is a useful risk reminder. Follow your own approved privacy and procurement rules.

What Not to Claim

Avoid claims such as:

Instead, say what the design does: “The task collects supervised planning, a working record and a short verification response.” Those are observable features.

Notify Students Precisely

State, in plain language:

Do not introduce a material restriction after students start. Align the notification with current NESA rules, sector requirements and the school's assessment policy. The 2026 HSC guide names unauthorised generative-AI use within its academic-integrity advice; see Before you start your HSC.

Preserve Access and Fairness

An oral verification step can disadvantage students if it becomes an unnotified second assessment or ignores approved adjustments. Keep it short, linked to the submitted mathematics and governed by a common protocol. Plan access arrangements through the school's existing process.

Likewise, requiring extensive handwritten drafts may measure speed, transcription or executive function more than the target outcome. Collect only the process evidence needed for a sound judgement.

A Faculty Review Checklist

Before approval, ask:

The best defence is not a one-off “AI-hard” question. It is a coherent assessment program that sees students solve, explain, revise and transfer mathematics across time. Pair this guide with writing an HSC maths assessment task and structuring a marking rubric.

Design for mathematical evidence

Build a question set you can review and adapt

Start from deterministic HSC maths questions, then add your own context, checkpoints and verification prompts.

Explore questions by topic

Related HSC resources

More from the blog