Flock has put an AI content moderator between police officers and its camera network. WIRED reconstructed the code the company ships to officers' browsers before login and found that a typed description of a person is checked against eight categories of sensitive content before it ever touches stored footage, and comes back with one of three verdicts: allow, block, or warn. The model runs on Flock's servers, so the department never sees the confidence score that produced the decision. And the category most heavily protected by the First Amendment — "political, social and cultural expression" — is the one category that is never blocked. It only warns, with a button underneath.
This is Flock's AI-powered search tool for police
Source: wired.com
The moderation layer exists because the rest of the product got hard to defend. Police departments in several US cities have torn up their Flock contracts, and city councils that had approved camera installations reversed themselves. Two members of Congress wrote to Flock CEO Garrett Langley after a sheriff's deputy in Texas ran a search across more than 83,000 cameras trying to locate a woman who had an abortion. In Illinois, the company gave federal immigration agents access to state camera data in violation of local law. Flock then announced a package of changes: a shorter default retention period, a case number required for every search, and an automated check meant to catch misuse. The company says these limits become mandatory by the end of the year.
WIRED examined the code behind those promises. Much of the search logic sits on Flock's servers and cannot be inspected from outside. What the client-side code shows is that the guardrails may talk some officers out of an abusive search and will preserve a record of it — but they do not stop the search from running.
The retrieval itself is ordinary machine learning. Every frame a camera captures is stored as a set of numbers describing its contents; a typed description is converted into numbers the same way; the model compares the two and surfaces the closest frames. Additional filters, thresholds and model instructions that can reshape the ranking are held server-side. A "Smart sort" feature lets an officer approve or reject individual images, after which the system reorders results toward whatever he approved. The interface suggests using it when there are too many hits.
Vehicle search is a menu: color, body type, make, model, and extra features such as a roof rack or bumper stickers. Person search has no filters at all — only free text. One of the examples built into the system is "person in medical scrubs." Flock calls this free-form search and says it has been selling it for years. 404 Media has reported that officers used free-form search to look for people by descriptions of their clothing and tattoos. Flock draws a line between the two: the vehicle search works on predefined characteristics, free-form works on natural language. On video cameras, free-form search covers clothing, colors, objects and other case-related details; on license plate readers it is restricted to vehicles, and person attributes cannot be searched.
The blocked categories are race or ethnicity, religion, nationality, "subjective or biased terms," and offensive content — sexual themes and innuendo, slurs and profanity, racist and other derogatory labels. In vehicle mode, anything describing a person is blocked, including gender, clothing and behavior. Religion and nationality are the only categories that can produce either a block or a warning the officer can click past. Political, social and cultural expression is described only as a warning category. Flock told WIRED that officers cannot search people by prohibited attributes and that attempts to use such terms will be blocked; it did not answer when the system warns instead.
Tom Bowman, policy counsel at the Center for Democracy and Technology's security and surveillance project, finds that split hard to justify: political and cultural expression sits among the most protected speech there is, and it is where police face the fewest restrictions. When a warning does fire — because of text on a shirt or a bumper sticker tied to protected speech — the officer is told the query will be logged and an administrator notified. He ticks a confirmation box, types a comment, and a button reading "Search anyway" becomes available. Flock justifies the escape hatch with legitimate cases: a victim describing a suspect in a biker gang jacket with a particular logo or emblem bearing a flag.
The taxonomy is also trivially routed around. The system bans searching by religion but not by clothing, so an officer can describe garments worn by members of a particular faith and get much the same result. Kate Ruane, who directs CDT's free expression project, points out that "political, social and cultural expression" is vague enough to cover finding out who attended a protest — which has already happened — and that the system treats the category as risky enough to warn about while handing over the results after a single click. No content moderation at this scale is fully accurate, she notes, and moving video is harder and more error-prone than text. One California officer typed "American flag." The person search blocked it. The vehicle search let it through and ran it across 11,000 cameras.
That flag query is the whole design problem in one line. Flock has built a filter that reasons about categories while the harm lives in intent, and it has put the filter in a place where no one can measure it. The model decides which descriptions of a human being an officer is permitted to look for, and nobody outside the company can check how often it is wrong or in which direction. Meanwhile the legal risk moves the other way: Flock warns officers that results may be inaccurate or incomplete and should not be the sole basis for a decision, while assigning "the risk of any inaccuracies" to the officer who ran the search. Responsibility flows to the user; the ability to see whether the system works stays with the vendor.
Notably absent from the company's answers is whether any of this binds. WIRED sent Flock a detailed account of its findings and 16 questions. The written reply said the tools should help police find information with clear limits, and that queries are checked against internal content rules. Flock did not explain how its model was built, what system instructions it runs on, or what separates one verdict from another. It did not say whether anything on its servers actually stops a query submitted without a reason or a case number — the code shows departments can currently configure some search types without either, and can switch the automated flagging off entirely. When the software does flag a search, Flock receives the query text, the category and confidence score the model assigned, and what the officer did next: continued, cancelled, or contested the warning. The company would not say how long it keeps those records, who inside Flock can read them, or whether they are used to audit or improve the moderation system. A firm selling accountability has made itself the least auditable node in the chain.
Deepak Kumar, an assistant professor of computer science and engineering at the University of California San Diego, reads the warning screen as a filter on resolve rather than on conduct: it stops officers who were not determined to begin with. He compares it to anti-harassment interfaces online, where motivated users sometimes do more damage after being warned, and to browser malware notices, where the effect depends on how the warning is presented. Logs can help administrators find officers who habitually click through, he says, but only where real oversight exists. As built, the system records the violation rather than preventing it. Screening the query before it runs does cut some abuse, Kumar adds, but without examining what the model returns it is not enough — checking both inputs and outputs is how most AI safety work is done now. Jay Stanley, senior policy analyst at the ACLU's speech, privacy and technology project, puts the filter in proportion: it is "a tiny detail on a huge machine of mass surveillance," a database an AI can be pointed at for fishing expeditions with almost no limits. Like Flock's other recent reforms, he argues, moderation targets individual bad officers and leaves the structural risk untouched.
The scale is the part the moderation debate obscures. These queries run against every person and vehicle the cameras have ever seen, as far back as a department keeps records. Besides text, the system accepts a license plate or a cropped photograph; it can show vehicles at locations an officer marks on a map, and surface travel-pattern associations — enter a plate and it returns cars repeatedly seen nearby. Reach depends on the mode, the department's permissions and its data-sharing agreements: its own cameras, cameras shared by other agencies, and for some plate queries, state and national networks. Two alerting systems run on top of this. One turns a text description into a standing watchlist across every camera in a map area the officer draws. The other, "People detection alert," needs no description at all: the officer boxes part of one camera's view — a doorway, a stretch of sidewalk — and is notified when a person enters it. One person is enough to trigger it. The system will not let the alert be set if confidence that the camera is seeing a person falls below 75 percent.
What that machinery has already been used for is documented. The Washington Post counted at least 50 US officers recently accused or charged with misusing license plate recognition systems; Flock appeared in 46 of those cases. In more than half, the target was a wife, a girlfriend, an ex-partner, an ex-partner's new boyfriend, or a woman the officer wanted to meet. A Milwaukee officer searched a woman he had dated 124 times and her ex-partner 55 times, logging each query as an "investigation." A Georgia police chief searched his ex-girlfriend and her daughter 600 times; he died by suicide five months after being charged. A Kansas police chief, never charged, ran 164 searches for an ex-partner and 64 for her boyfriend, and lost his job. A North Carolina officer searched his boyfriend's ex-wife 31 times and filed 29 of the queries as traffic violations. None of this began with Flock: an Associated Press investigation found officers mining police databases for ex-partners and people they found attractive, with more than 325 officers fired, suspended or pushed to resign over a two-year stretch across state agencies and dozens of the largest US police departments. WIRED has reported at least 414 investigations into Immigration and Customs Enforcement employees and contractors for misusing sensitive government databases, and has since found hundreds more allegations against Customs and Border Protection staff, including officers tracking women they wanted to date, monitoring family members and following colleagues' phones.
Flock's answer is that much of the responsibility for usage rules belongs to local departments, which set access and procedure. Don De Lucca, former police chief of Miami Beach and past president of the International Association of Chiefs of Police, broadly agrees — vendors can recommend safeguards, but the rules have to be written before the technology is deployed, and access should be tied to a lawful basis. A special investigations commander at a Southern California department, who asked not to be named because he is not authorized to speak publicly, argues the exposure is not specific to Flock or to surveillance: new tools routinely create room for both legitimate use and abuse, and what matters is whether the deployment comes with hard limits, continuous oversight, auditability and consequences.
Which is where Flock's own numbers turn against the narrative. As pressure mounted, the company added an audit-assist feature that flags suspicious patterns worth reviewing: the same officer repeatedly querying one plate, searches confined to another department's cameras, one plate searched under several case numbers. Flock says more than a third of its customers have switched it on — and in the code WIRED examined, the feature is still off by default. A South Carolina sheriff's office turned it on recently. The next day, an internal security officer found more than 2,700 apparently unauthorized queries.
One office, one day, 2,700 queries. That number says the constraint was never officer behavior; it was detection, and detection had been optional. Flock's fix is to make it mandatory by the end of the year, which will convert an unknown volume of misuse into a measured one — and the measurement will be generated, categorized, scored and retained by the company whose product is being measured, on servers no department can see into.