What the record is — and isn't
Does Phloem mark or grade student work?
No — and it never will. Phloem doesn't score writing, grade it, or judge its quality. The report describes how a document came to be: which sentences were typed, which were pasted in, which came from an AI suggestion taken whole — and which came from a suggestion the writer read, thought about, and rewrote in her own words. What a teacher makes of that description is the teacher's call, within the context of everything they already know about the student.
The Ministry of Education's guidance on AI and marking puts its first principle this way: "AI must support — not replace — teachers' professional judgements." That principle is Phloem's design premise, not a constraint it works around. The report gives a teacher more to exercise judgement on, and takes none of it away.
Marking work with AI tools, Ministry of Education / Te Poutāhū Curriculum Centre, October 2025 (PDF). The principle is printed there verbatim.
Is this an AI detector? Why not just use one?
No. A detector reads a finished text and guesses where it came from. Phloem doesn't guess, because it doesn't need to: the record was kept while the writing happened, event by event, in a sealed log the writer owns.
The guessing is getting better, and it is still the wrong instrument. In a study by Perkins and colleagues (samples tested September–October 2023, published 2024), seven major detectors averaged 39.5% accuracy on unmodified AI-generated text and 17.4% once the text was lightly disguised; human-written control samples were correctly identified only 67% of the time, and the authors put the rate of false accusations at 15%. Two 2026 studies in the same journal show how far things have moved — and how far they haven't. Van Vlasselaer and colleagues (Vrije Universiteit Brussel, June 2026) ran four detectors over 160 long academic papers with known origins: all four tools placed every one of the forty human-written papers in the clear band, with zero false positives for three of them — false positives, the old fear, were "almost absent" — while Turnitin, the tool most institutions actually run, identified none of the forty fully AI-written papers as AI. That figure needs its mechanism, in fairness to Turnitin: it publishes no percentage at all between 1 and 19 per cent, showing an asterisk in place of a number, because suppressing that band is how it keeps false accusations rare — and nineteen of those forty papers came back as exactly that asterisk, an admission of uncertainty rather than a confident wrong answer. Turnitin's own chief product officer states the trade openly: the company is "comfortable sacrificing accuracy in identifying AI-generated writing in order to deliver a less than 1 percent document false positive rate." Designed, then, rather than broken — and the consequence is the same either way: the tool is tuned to stay quiet, and quiet is not the same as knowing. Only one newer tool, Pangram, reliably caught AI text that had been "humanised". Hadra and colleagues (February 2026) tested Turnitin and Originality.ai — an independent detector, not to be confused with Turnitin's own product of a similar name — on hybrid texts — human and AI mixed, "an increasingly common form of student writing" — and found both weak exactly there: of the hybrid texts, Turnitin correctly identified 31% and Originality 2%. Accuracy also fell sharply on scientific writing, and Originality showed a borderline tendency to misread students writing in English as a foreign language. The picture isn't uniform — the Brussels study rated Turnitin the best of its four tools on hybrid text, at 60% — and that disagreement is itself the finding: the same tool, on the same category of writing, scores very differently depending on whose test it faces. Which is the problem with evidence you cannot check. Both papers reach the same conclusion: a detector can be an initial flag, never the evidence. Every study above is dated in the sentence that cites it, and every one is linked below, so you can read them yourself rather than take our word for the summary.
That is the point a better detector doesn't fix. A score on a finished text cannot be checked in the field, because nobody knows where the text truly came from; and it cannot see work done with AI rather than by it — which, as Bassett and colleagues observed in January 2026, is how students actually write now, making the human-or-AI binary "not merely inadequate but meaningless".
There is also a reason the numbers can't rescue it, and it is worth understanding because it survives every improvement the detectors will ever make. Suppose a detector publishes a false-positive rate of 1% — one human paper in a hundred wrongly flagged. It is tempting to read a flag as therefore 99% likely to be right. That inference is simply invalid. To know what a flag means you need three things: the false-positive rate, the true-positive rate, and the base rate — the actual proportion of AI-written work in that particular cohort. Nobody knows that proportion. Nobody can: measuring it would require knowing the truth about each paper, which is the very thing in question. So the flag's real meaning cannot be computed, and a percentage that cannot be interpreted cannot support a finding on the balance of probabilities. This is not an argument that today's detectors are poor. It is an argument that no score of this kind, however accurate its maker, can carry the weight a fair process requires.
The honest writer accused by a percentage has no way to answer, however good the percentage. A writer with a Phloem record does: she hands it over, sealed, and says — here's how it came to be.
None of this replaces your plagiarism checker. It runs over the finished text exactly as it did yesterday; Phloem neither replaces it nor obstructs it.
Perkins, Roe et al., Simple techniques to bypass GenAI text detectors, International Journal of Educational Technology in Higher Education 21:53 (2024) — the published version of the study also circulated as a preprint; its abstract says six detectors and its methods say seven, and the figures quoted here are from its results · Van Vlasselaer, Van Droogenbroeck & Spruyt, Who wrote this? Evaluating the reliability of AI detection tools in higher education, International Journal for Educational Integrity 22:16 (29 June 2026) · Hadra, Cambridge & Mesbah, Evaluating the accuracy and reliability of AI content detectors in academic contexts, IJEI 22:4 (2 February 2026) · Bassett et al., Heads we win, tails you lose: AI detectors in education, Journal of Higher Education Policy and Management (29 January 2026) · Liang et al., GPT detectors are biased against non-native English writers (2023; the 2026 studies above found far fewer false positives on human writing than Liang did) · Annie Chechitelli, Chief Product Officer, Turnitin, letter to The Chronicle of Higher Education (August 2026) · Turnitin's asterisk convention is documented in its own AI Writing Report guide.
Isn't this just surveillance with better manners?
Opposite polarity. Surveillance takes evidence from the student; this is evidence she offers. The record is hers — shared by her consent. She keeps her own copy.
Although, it is fair to say that the moment an assignment brief says "attach your Phloem report", consent becomes compliance. We don't pretend otherwise. What survives that moment is still hers — her own copy of the record, and a description rather than a score.
Couldn't a student pass copied work off as her own typing?
Yes — by retyping it, and the report will not catch her. The report never uses the word "original", because it can't know: it records typed, pasted, offered-and-taken, offered-and-declined — facts about how the document came to be, never certificates of where a sentence was born. I found the limit myself, writing a test essay with Wikipedia open in the next window: my transcription went down as my own typing, because it was my own typing. No tool can see the second screen — including the keystroke-watchers whose scores imply they can.
What the record does close off is the cheap version. A paste wears its own mark. AI help is on the ledger along with what she did about it. The deception that remains costs her hours of reading and retyping at human speed — which is, not coincidentally, the very time-on-task the assignment was asking for.
But she could type it all herself and learn nothing.
So could I, forty years ago — cram for three days, write the essay, lose ninety-five percent of it within the week. The finished paper has never proved learning; it only ever proved its own existence. What's changed is where the facts live: a date or a formula is seconds away, so education has moved to what it was always really about — how to think, how to learn, how to work with sources.
A truthful account of the process serves that; what the margin offered, what she took, what she declined, what she reworked in her own words. It isn't proof of learning — nothing is — but it is more than the finished paper ever told you.
The writing, the law, and the machine
Is students' writing used to train the AI?
No. The margin runs on Anthropic's Claude through their commercial API, and the Commercial Terms say it in one clause: "Anthropic may not train models on Customer Content from Services." We can put that clause in front of you.
The Ministry of Education's own guidance warns about precisely this: many AI tools reuse what's typed into them as training data, and putting student work into a tool that isn't information-protected "may be in breach of NZ Privacy Law" (October 2025). That warning is the right test to apply — to Phloem as much as to anything else. It is the test Phloem was built to pass: the terms forbid training on the writing, and the data-processing agreement that governs it is a document we can hand you, not a reassurance.
Anthropic Commercial Terms of Service §B and the Data Processing Addendum incorporated into them (§B effective 17 June 2025, clause re-checked 25 August 2026) · Marking work with AI tools, Ministry of Education, October 2025 (PDF).
The writing goes to a model in the United States. Doesn't that need consent?
When the writer summons the margin, the relevant writing goes to Anthropic's servers in the United States, under the terms above.
In law, that is processing, not disclosure. Anthropic processes the information on the school's behalf, so under section 11 of the Privacy Act 2020 the school never stops holding it — which makes this a security question under IPP 5, not a cross-border disclosure under IPP 12. That isn't our reading alone: it's the Privacy Commissioner's published position.
Privacy Act 2020, s 11 · OPC — Sending information overseas: an agency using a cloud provider remains responsible for the information; holding under IPP 5, not disclosure under IPP 12.
If something goes wrong, is that on you or on the school?
On the school — the information never leaves the school's hands legally, the school carries the responsibility, including breach notification. Which is exactly why a school should hold its providers to a contract that makes that safe, rather than to a reassurance.
Can a student have her record deleted?
Anything a student writes into Phloem belongs to the student, and stays with the student until she chooses to share a report. There is no copy held anywhere else for her to ask deleted. When she does choose to share one, the copy she hands over is the school's to hold, and any privacy obligations that come with it rest with the school from that moment — governed by the school's own rules, not ours.
What if a student writes about somebody else — a sibling, a teacher?
That's IPP 3A, which came into force on 1 May 2026 — new law about collecting personal information indirectly. Our reading is that the exceptions carry most of the school writing case, but it is new law and our reading is marked as exactly that: a school's privacy officer should reach their own view rather than take ours.
Privacy Act 2020, IPP 3A, in force 1 May 2026 · OPC — the privacy principles in plain language.
Running it in a school
Do we have to install anything? Does IT need to approve it?
No. The student writes, exports her report — one file, the readable report carrying its own sealed record — and attaches it to the assignment she already submits. Nothing installed, no tenant administration, no procurement, no new account for the school to govern.
Are you a Ministry-approved vendor?
No, and we're not asking to be. The Ministry holds cloud agreements with Microsoft and Google; Phloem isn't on that list, and that is precisely why the pilot attaches a file to an assignment you already run instead of asking the school to adopt a new cloud service holding student writing. The report travels inside the systems you've already approved.
Has a lawyer reviewed any of this?
No. Everything on this page is our own reading of published guidance and primary sources, written up with the gaps marked in the documents themselves. We'd rather hand you that than a confident answer nobody has checked. The two working documents behind these answers — the Privacy Act analysis and the school-as-agency analysis — exist, are dated, and mark what we don't know; if fifteen minutes with them would help your privacy officer, write and we'll send them.
What's your data-breach process?
Owed, not built. Anthropic is contractually obliged to notify us of a breach; nothing yet obliges us to notify a school. Meanwhile it is the school, as the agency holding the information, that carries the legal duty to report a serious privacy breach to the Privacy Commissioner and the people affected — and the Commissioner's published expectation is notification within 72 hours of learning of it (the Act's own test is "as soon as practicable"). A breach-notification agreement is a short document, and it's the first one we'd sign before a single student writes under a school's roof.
How it compares, and what it runs on
How is this different from Turnitin Clarity or Grammarly Authorship? They record the writing process too.
They do. Three differences, all structural. Who holds the record: theirs lives on the vendor's servers or inside the institution's learning system; Phloem's is the writer's own file, and it verifies without us. What it knows: they record what entered the editor; Phloem also knows what the writer was shown — every AI suggestion is an offer on the ledger, so the report can tell taken-whole from edited from reworked in her own words. What comes out: both put a statistical AI score alongside the record — Grammarly's inside the Authorship report, Turnitin's inside the Writing Report wherever the institution also licenses its detector; Phloem refuses to. It describes, and stops.
The honest other half: they are shipping products — Grammarly counted more than a million Authorship reports in its first year — with integrations into the systems institutions already run. Phloem is a working prototype with one writer. Same genus, opposite polarity — and a long way behind on everything except the polarity.
Grammarly Authorship (2024–) and Turnitin Clarity (GA July 2025), as read August 2026 from their own product pages; the AI-score behaviour is documented in Turnitin's own Clarity guide. The Chronicle of Higher Education, 5 August 2026, on Clarity.
Does it run on Chromebooks or Windows? Does it work inside Word or Google Docs? Does a student need an account?
It runs in any modern browser — that is where it was built — and as a desktop app on the Mac. It is its own writing surface, not a plug-in: the record is the document, so it can't be bolted onto Word or Docs. No student account, no API key, no sign-up ritual: the hosted version opens from a link, and the AI usage is covered on our side within sensible daily limits. The browser version carries a little less than the desktop one — no live folder of source documents, and exports arrive as downloads — and there is no Windows desktop app yet.
What's the evidence it works? How many students have used it?
None on graded work — that is why this page exists before a sales page does. What has been tested: the report's reading of the log was scored against a hand-kept record of a real essay, thirteen paragraphs out of thirteen; it has since been hardened against sealed test documents and a second, independent implementation that has to agree with it; dictation detection failed its first real test and was rewritten until it passed; and the sealed record was attacked before it was trusted. One writer, one evening of outside testers, and a demonstration report you can read. The next thing it needs is a classroom — fifteen minutes of real writing, then the report, then tell us what it says about how you wrote. If you can offer that, write.
What we haven't done yet
- No practitioner has reviewed the legal analysis. It is our own reading of published guidance.
- No breach process and no school agreement exist yet. Both are short documents, both owed before any student writes in a school pilot.
- The sources are dated, deliberately. Terms of service and government guidance change without announcing themselves; every claim above says when we read it, so it can be checked rather than trusted.