The first category is manual paraphrasing. A student generates a draft with a language model, then rewrites it sentence by sentence, sometimes pulling vocabulary or sentence structure from prior coursework to match their voice. This is the oldest technique and remains the most effective, because the final text is in a real sense human-written even if the ideas and structure are not.
The short version
Students bypass AI detection through three main strategies: manual paraphrasing, commercial humanizer tools, and mixed human-AI authorship. Peer-reviewed research, including the DIPPER paper from 2023 and adversarial paraphrasing studies from 2025, shows that paraphrasing degrades detector accuracy substantially. Commercial humanizers including Walter Writes, Undetectable.ai, and StealthGPT operate on similar principles. Detection-only strategies are fragile; assignments that incorporate process evidence and oral defense are not.
What follows treats bypass as an engineering problem rather than a moral one. Students who reach for these tools are responding rationally to assignments where the cost-benefit of writing the thing themselves has tilted in the wrong direction. That is a design problem before it is a discipline problem.

The Three Categories of Bypass
The second category is commercial humanizer tools. These are services that ingest AI-generated text and return a version intended to register as human on detection software. The market includes Walter Writes, Undetectable.ai, and StealthGPT, among others. They share a basic architecture: a paraphrasing model trained or tuned to disrupt the statistical signatures that detectors rely on. They differ in price, interface, and aggressiveness of rewriting.
The third category is mixed authorship. A student writes part of the assignment themselves, prompts a model to write another part, and stitches the result together with light editing. This is harder to characterize as cheating in a binary sense, and it is also harder for detectors to identify reliably, because the resulting text contains both human and machine signals interleaved.
A paper that registers as human on a detector can still fail an oral examination, a process audit, or any assessment that asks the student to engage with the content rather than the artifact.
Working Educators editorial position
What the Research Shows About Paraphrase Attacks
The seminal paper here is Krishna et al.'s 2023 work on DIPPER, a paraphrasing model designed to test the robustness of AI text detectors. The finding, in plain terms, was that running AI-generated text through a strong paraphraser dropped detection accuracy substantially across every detector the authors tested. The paper is worth reading in full because it grounds a debate that often gets conducted in vibes.
More recent work has extended these findings. A 2025 paper on adversarial paraphrasing tested detector performance against paraphrasers that were explicitly optimized to evade detection, and found that even leading detectors degraded significantly under these conditions. The relevant point for teachers is not that detectors are useless, but that any pedagogy that depends on detectors as a single point of truth is building on sand.
Our earlier piece on the Stanford TOEFL study covers an adjacent finding, which is that detectors can produce systematic false positives on non-native English writers even without any bypass attempt. The combination of false-positive risk and adversarial bypass risk is what makes detection-only enforcement fragile.
Commercial Humanizers and What They Actually Do
The commercial humanizer market emerged because the paraphrasing finding became public knowledge and someone built a product around it. The three tools we named above are the most visible, but the category has dozens of entrants and the underlying technology is converging.
What these tools do, technically, is paraphrase. They take input text and produce output text intended to disrupt the perplexity, burstiness, and stylistic markers that detectors use to classify writing as machine-generated. Some operate as single-pass rewrites; others iterate, scoring their own output against a target detector and re-running until they hit a threshold. The marketing copy on these sites varies in its claims, and we have not seen independent third-party testing that would let us rank them on bypass success rates. We mention them by name because their existence is part of the landscape any teacher should know about, not because we endorse them. Walter Writes, Undetectable.ai, and StealthGPT are reasonable representatives of the category for purposes of understanding what students may have access to.
One observation worth making: these tools are priced for individual student budgets, available without institutional gatekeeping, and improve quickly. Pretending they do not exist or treating them as exotic is unhelpful. Treating them as a permanent feature of the writing environment is more useful, because it forces the question of what assignment design looks like when this kind of tool is universally available.
Why Most Bypasses Fail at Higher Stakes
Detection at the document level is one thing. Defending the document is another. A paper that registers as human on a detector can still fail an oral examination, a process audit, a follow-up question, or any assessment that asks the student to engage with the content rather than the artifact. This is the central insight that detection-only conversations miss.
Consider what a student who used a humanizer cannot easily do. They cannot describe the dead ends they hit in the research process, because there were none. They cannot point to a specific passage and explain why they made the rhetorical choice they made, because they did not make it. They cannot reconstruct the argument from memory without referring back to the document, because the argument was never theirs to begin with. None of this requires a detector to surface; it requires a conversation.
This is why the most durable answer to bypass is not a better detector. It is an assessment structure that produces multiple windows onto the student's actual understanding. Detection becomes one data point among several rather than the load-bearing element of academic integrity. We discuss this in more detail in our piece on Turnitin AI scores and academic integrity.
What Works Better Than Detection Alone
The pedagogical literature has converged on a small set of practices that hold up well against bypass, and that also happen to be good teaching independent of AI. The first is oral defense or follow-up conferencing, in which the student walks an instructor or TA through their submission and answers questions about it. Our piece on oral defense as assessment goes deeper on this, but the short version is that oral defense is robust to every form of bypass we have discussed.
The second is process-based assessment, in which the student submits evidence of the writing process itself: drafts, notes, outlines, version histories, annotated readings. A 2025 Frontiers in Computer Science paper on process-based assessment in the AI era provides a useful framework. The idea is not new; what is new is the renewed reason to take it seriously.
The third is assignment redesign that makes generative AI a less useful tool, either by tying the assignment to material the student worked with in class, by requiring the integration of specific local context, or by asking for analysis of student-generated artifacts that the model cannot have seen. An Inside Higher Ed essay from April 2026 argues this point well, and we agree with its core claim that the best defense against AI cheating is teaching that is harder to phone in.
None of these alternatives are free. They all cost faculty time, and the time has to come from somewhere. Our editorial view is that the time spent designing bypass-resistant assignments is more durable than the time spent adjudicating detector flags, and that institutions that recognize this will produce better outcomes for students and faculty alike. Our piece on redesigning the research paper covers a concrete example of what this looks like in practice.
Frequently Asked Questions
Should I be running every submission through a detector and a humanizer to see what comes back?▼
No. That kind of arms-race testing produces noisy results, consumes time that would be better spent on assignment design, and trains you to chase a moving target. The more productive question is whether your assessment structure depends on detection at all, and if so, what would happen if detection stopped working.
Are commercial humanizers illegal or against terms of service?▼
They are legal products in most jurisdictions. Whether their use violates your institution's academic integrity policy is a separate question, and the answer depends on how your policy is written. Most policies treat undisclosed AI use as a violation regardless of whether a humanizer was used; the humanizer is a means, not the offense.
What about students who only use AI for brainstorming and then write everything themselves?▼
This is the case where blanket prohibition is most clearly counterproductive. A policy that distinguishes between generative use and assistive use, and that requires disclosure of either, handles this case without forcing teachers to adjudicate the impossible question of where one stops and the other begins.
If detection is unreliable, why use it at all?▼
Detection can be useful as one signal among several, particularly when paired with process evidence and oral follow-up. The failure mode is treating it as conclusive on its own. A flag is a reason to ask questions, not a verdict.
Will humanizers eventually become undetectable?▼
For practical purposes, in some classroom settings, they already are. This is one of the reasons we argue that pedagogy that does not depend on detection is the more durable answer. The arms race between humanizers and detectors will continue, but the assignment that asks a student to defend their own argument in person is robust to its outcome.
The Bottom Line
Bypass tools exist, they are inexpensive, they are improving, and treating them as a moral problem rather than a structural one will not make them go away. Students who reach for them are responding to assignments where the path of least resistance has tilted away from doing the work, and that is a design signal worth heeding.
Our editorial position is that the most useful frame for teachers is to ask what their assessment would look like if detection stopped working entirely. If the answer is that nothing would change, the assessment is robust. If the answer is that the entire integrity model collapses, the assessment needs work. Most of the alternatives are not new; they are practices that good teachers used before detectors existed and will continue to use after detectors stop being useful.
We do not endorse any of the humanizer products named in this piece, nor do we condemn them. They are part of the writing environment students are working in, and pretending otherwise serves no one. The more honest conversation is about what good assessment looks like in a world where they exist.


