Research story · Education, AI & procedural fairness
What did the
tool contribute?
An assessment should establish what intellectual work the student performed. Published research examines how universities can distinguish faithful format conversion from machine contributions to the substance of a submission.
A student completes a mathematical derivation by hand, then uses an AI-powered tool to turn the page into typeset text. Another student asks the same platform to produce the derivation. The finished documents may look similar, but the direction of the intellectual work is different. A sound assessment policy needs a way to establish that difference.
My article, Transcription is not generation: distinguishing non-generative AI tool use from academic misconduct in higher education assessment, published in the International Journal for Educational Integrity, develops a function-based account of the problem. It connects technical classification with assessment validity and a proposed framework for evaluating evidence.
Begin with the learning outcome
The central question is what the assessment is designed to measure. If it tests a student’s reasoning, analysis or proof, faithful conversion of an already-authored answer may leave that intellectual contribution intact. A tool that supplies the reasoning changes the contribution itself.
If the assessment instead tests transcription, symbol recognition or a note-taking process, using a tool to perform that operation can bypass the skill under assessment even when no new argument is generated. The article makes this boundary explicit. Its interpretation does not grant every transcription task an automatic exemption.
The applicable policy also matters. A rule expressed in terms of generative function is different from an explicit prohibition on a named platform or specified tool. The paper defends an interpretation and a policy-design proposal; it does not claim that every institution currently uses identical wording or permits the same assistance.
Describe the operation rather than the brand
Modern platforms combine many capabilities. The same interface can recognise handwriting, convert speech, suggest revisions or generate a complete answer. Knowing that a platform was used establishes neither which operation occurred nor what the student accepted into the final work.
The manuscript separates three categories. Novel content production creates assessed intellectual material. Error-correcting reconstruction occupies a boundary where recognition can shade into inferred or altered content. Format conversion preserves the substance of material already authored by the student.
The middle category is important. A request labelled “transcription” can still produce an unsolicited correction or explanatory addition. Conversely, use of a model capable of generation does not establish that generation supplied the assessed reasoning in a particular interaction. The evidence must address the actual function and resulting content.
Make the distinction examinable
A conceptual boundary is useful only if a decision-maker can examine a disputed case. The article proposes four mutually supporting criteria: fidelity, non-augmentation, traceability and attestation. Together they shift the inquiry toward the relationship between original work, tool interaction and submitted output.
Fidelity asks whether the typeset version corresponds to the pre-existing material in substance. The same equations, reasoning steps and conclusions should remain identifiable. Recognition errors and formatting differences require careful assessment; a correction to the mathematical reasoning is a different kind of change.
Non-augmentation asks whether the tool supplied substantive suggestions, missing steps or alternative approaches. The inquiry concerns what was contributed, rather than whether the output looks polished. It includes the possibility that a student accepted material the tool offered without being explicitly asked.
Traceability follows the workflow through contemporaneous originals, submitted images, instructions and intermediate outputs. Attestation records the student’s account of checking fidelity and excluding substantive additions. The framework treats those materials as evidence to scrutinise, not as an infallible certificate.
Keep the responsibility for judgment clear
The proposed criteria give a student a structured way to explain a transcription-only workflow. The institution retains responsibility for evaluating the evidence and establishing a breach under its applicable rules. Producing records does not automatically settle the case, and their absence does not turn an unvalidated stylistic impression into proof.
Logs can be incomplete or selectively presented. Originals and outputs need to be considered together, and an account of the process can require corroboration. The criteria verify the transcription step, not authorship of the handwritten original: that original could itself reproduce material generated elsewhere. Establishing a faithful conversion therefore does not by itself establish who performed the upstream intellectual work.
Its empirical questions are therefore practical: can non-specialist panels apply the criteria consistently, do the criteria distinguish cases accurately, and what does implementation require? The article proposes that work; it does not report a validated instrument or a study showing the framework’s performance in misconduct hearings.
Separate a signal from a finding
The article reviews research on AI detection and stylistic heuristics, including findings in both directions. Some studies report useful detection signals under their tested conditions. They also document limits, variation across text types and the difference between detecting some machine-generated content and reliably establishing what occurred in a particular submission.
The paper’s concern is evidential sufficiency. A detector output is an indirect signal about text characteristics. A misconduct decision concerns prohibited conduct, authorship and the governing assessment rule. Those are connected questions, but they cannot simply be collapsed into one score.
The same applies to impressions of polished, structured or consistent prose. Such features can arise from disciplinary training, typesetting or faithful transcription as well as generation. The research argues for examining plausible explanations against workflow evidence instead of treating appearance as a unique signature of misconduct.
Design fair rules that protect genuine assessment
Accessibility adds another reason for precision. A tool can support the production of legible output while leaving intellectual authorship with the student. Policy design needs to address that function alongside the assessed learning outcome and the institution’s adjustment arrangements, rather than treating every assistive use as interchangeable with generated work.
The proposal also protects the purpose of genuine restrictions. A clear distinction between format conversion and intellectual augmentation makes it easier to explain why one use may preserve an assessment while another undermines it. Functional specificity can strengthen the rule by making its target explicit.
This is a conceptual and policy contribution bringing computer science, education and procedural fairness into one framework. It does not estimate the prevalence of wrongful accusations or establish a universal detector failure rate. It identifies the distinctions a defensible evaluation needs to preserve.
The practical research agenda is to make authorship examinable: identify the learning outcome, state the permitted functions, reconstruct the actual workflow and assess the evidence with appropriate care. That approach keeps academic integrity focused on the intellectual contribution the assessment exists to measure.