The Control Journal
GuidesSeptember 16, 20267 min read

An AI Proctor Joined Your Interview: What It Can See

Live-interview proctoring agents join Zoom, Meet, and Teams as participants. What that architecture can observe, what it cannot, and how to tell.

CControl Editorial Team

A new class of hiring tool joins your video interview as a participant and watches you for signs of fraud. Vendors call these interview proctoring agents. Architecturally they are meeting bots, and that architecture sets a hard limit on what they can observe: the audio, video, screen share, and chat your machine sends into the meeting, plus meeting metadata. Anything beyond that boundary — your keystrokes, your mouse, a second monitor, the apps open behind your call window — requires software running on your computer or inside your browser, not a participant in the call.

That distinction matters because several vendors advertise both kinds of capability on the same page without saying which product configuration delivers which. This article separates the two, then explains how to tell, before an interview starts, which one you are actually facing.

What a live-interview proctoring agent is

An interview proctoring agent is a service that attaches to a scheduled video interview and analyzes the call in real time for indicators of cheating, impersonation, or AI assistance, returning alerts to the interviewer and a report afterward.

Sherlock AI markets itself as an "Interview Proctoring Agent" that works with Zoom, Google Meet, and Teams, connects to Google, Apple, or Outlook calendars, and is toggled on per meeting. As of September 16, 2026, its site describes "a multimodal adversarial ML approach to detect interview fraud, combining signals from device activity, audio environments, and candidate behavior into a unified classifier," and reports raising "overall detection accuracy from ~85% to over 97%" (sherlock.sh). That figure is a vendor self-report; no methodology, dataset, or independent evaluation accompanies it.

CheqMate AI describes a similar calendar-connected workflow across Teams, Google Meet, and Zoom, and claims it "Flags 100+ AI interview tools — from hidden overlays to background copilots," along with detection of "Hidden screen overlays, second screens, and covert coaching setups" and keystroke and mouse pattern analysis for "remote desktop takeovers" (cheqmateai.com). Both are vendor claims, not independently verified results.

These are not the same thing as an AI notetaker. A notetaker transcribes and summarizes; a proctoring agent scores you for integrity risk. If you are trying to work out which bot is in your call and what obligations come with it, the rules covering AI notetakers in job interviews cover the transcription side of the question.

The capture boundary a meeting bot cannot cross

A meeting bot is, in the words of Recall.ai — the infrastructure vendor many of these products build on — "a virtual participant that joins video calls (Zoom, Google Meet, Microsoft Teams, Slack Huddles, etc.) to ingest meeting data and sometimes output data back into the meeting in real time" (recall.ai).

What a participant-level bot can access is exactly what any attendee's client receives:

  • meeting audio in real time;
  • participant video tiles and any screen share;
  • messages posted in the meeting chat;
  • join and leave events, mute state, and camera on/off state;
  • meeting metadata such as title, host, participant emails, and duration.

Recall.ai's documentation also notes that these bots "are always visible to everyone in the meeting, which may not be desirable for every workflow." Visibility is a property of the platform, not a courtesy.

What that list does not include is anything happening on your machine outside the meeting. A bot sitting in the participant roster has no view of your window list, your process table, your clipboard, your second display, or the timing of your keystrokes. It sees the pixels you chose to share and the video your webcam sends. If you share one window rather than a full desktop, the bot sees one window — the same constraint that governs what an interviewer sees during a screen share.

Which vendor claims fit inside the boundary

Sorting the marketing by what the architecture supports is the fastest way to read one of these product pages.

Plausible from a meeting stream alone

Face matching against an ID photo, liveness and deepfake artifact analysis, gaze direction estimated from the webcam tile, speech cadence and reading-aloud patterns, background audio analysis, and detection of a second voice are all derivable from audio and video that the bot legitimately receives. Whether any given vendor does them well is a separate question, but the inputs exist.

Not derivable from a meeting stream

Keystroke timing, mouse movement patterns, remote desktop takeover, running-process inspection, and "second screen" detection are not available to a participant. Those signals require an agent installed on the candidate's machine, a lockdown browser, or a browser-based assessment page with instrumentation — a different product surface with a different consent story. The same claim appears in assessment-platform marketing, where monitor detection is usually inferred from browser APIs rather than observed directly; the mechanics and their failure modes are covered in how coding interviews infer a second monitor.

When a vendor lists both categories under one heading, the honest reading is that the product has more than one deployment mode. Ask which one your interview uses.

How to tell what you are facing

Before the interview, and in the first minute of it, four things are usually observable.

The participant roster. A meeting bot appears there, typically under a product or assistant name. If an unfamiliar named participant joins and never speaks, that is your proctoring or notetaking layer.

The recording disclaimer. Zoom documents that it "will always notify meeting participants that a meeting is being recorded," with an on-screen consent disclaimer for app users and an audio prompt for phone participants; on Enterprise and Education accounts, admins can disable the disclaimer for internal participants but "The recording consent disclaimer is required for all guest participants" (Zoom support). As a candidate you are a guest, so you should see it.

A consent prompt on Google Meet. Google Workspace administrators can "require attendees to give their consent before certain features are used, including transcription, recording, and AI note-taking," and can decide whether declining removes you from the meeting or stops the feature. That setting applies to mobile and web users; participants joining on Meet hardware or by phone "are automatically consented" (Google Workspace admin help).

Anything you are asked to install or permit. A download, an extension, a "secure browser," or a browser prompt for camera, microphone, and screen-capture permission on the assessment page means the monitoring is running on your side of the connection, not in the call. That is the configuration where keystroke and display claims become technically possible.

Limits, error rates, and what a flag is worth

Accuracy claims in this category are self-published. A move from roughly 85% to over 97% detection, as Sherlock states, is meaningful only alongside a false-positive rate, a defined population, and a description of what counts as a positive — none of which these pages provide. Behavioral inference from a webcam tile is noisy by nature: looking away, reading a question twice, a rehearsed answer, an accent the model handles poorly, or a slow connection can all move a score.

Treat an integrity flag as the start of a process rather than a verdict. Most platforms surface a risk indicator to a human reviewer rather than automatically rejecting anyone, and the reviewer's interpretation is where the outcome is decided — the sequence is laid out in what happens after a coding assessment flags you. If a flag is raised against you, ask what specific signal produced it and what evidence was retained.

Where this leaves a candidate

Read the invitation before the call. If the interview is a plain Zoom, Meet, or Teams link with a bot in the roster, the observable surface is your camera, your microphone, and whatever you share. If the invitation routes you to a proctored assessment page or asks you to install something, assume machine-level monitoring and read the policy you are agreeing to.

Control is a desktop AI interview assistant for interviews, assessments, and screen-share workflows. As of September 16, 2026, its site states the overlay does not appear in the window you share and lists scope for that behavior — "Invisible to Zoom (≤6.16) and any browser-based screen recording software" — which is a claim about the shared-window surface, tied to specific versions, and not a claim about webcam-facing analysis, audio analysis, or any employer's policy. A tool that stays out of a screen share does not change what a camera pointed at your face records, and it does not resolve whether the role you are applying for permits assistance. Those remain your decisions to make against the rules you were given.

If you want to see how a desktop assistant behaves against your own screen-share setup before an interview rather than during one, the free tier on the pricing page includes 5 messages and 2 minutes of voice, which is enough to check the behavior on your own machine.

Continue exploring

Control AI - An AI Proctor Joined Your Interview: What It Can See