The longer story
AI detectors produce both false positives and false negatives, and they have repeatedly been shown to flag work by students writing in a second language more often. No defensible process treats a detector score as proof. What tends to be more telling is the distance between a submission and what you already know of that student: a sudden shift in register, references that turn out not to exist, or an argument that never lands on a specific detail from your lessons. The most robust responses are procedural rather than forensic, such as asking the student to talk through their reasoning, or building in drafting stages you can see along the way.

