Big MLSs know that AI is getting their data—either from careless agents or through “scraping.” How are they addressing the risks?
Bright MLS, a major U.S. MLS with over 100,000 subscribers, is actively engaging with vendors regarding AI model training on its listing data. They have developed a licensing structure that includes AI training as a specific addendum to standard data agreements. By mid-September, Bright MLS plans to launch a Model Context Protocol (MCP) server. This server will enable subscribers to directly link an AI assistant to Bright’s data, eliminating the need to download raw files or paste information into external tools, thus ensuring compliant AI usage by brokers and agents. The organization believes in an API-based metering system where data is called when needed, allowing for usage tracking and governance, a model smaller MLSs may adopt.
Data governance poses a significant challenge for MLS leaders like Bright MLS. Rajeev Sajja of Bright MLS confirmed that major AI platforms were using copyrighted listing photos and descriptions from their listings without authorization. Although Bright successfully halted this unauthorized use for now, Sajja is skeptical about future voluntary compliance, noting that AI companies often seek forgiveness rather than permission. Art Carter of California Regional MLS (CRMLS) expressed similar concerns. He tested a chatbot, Claude, which suggested either direct licensing or using a browser extension to scrape data from a subscriber's account. Both methods are problematic, as unauthorized plugins bypass MLS oversight. CRMLS is addressing this with NexusRE, a governance layer that monitors and controls how AI platforms access and monetize listing data, enforcing permission-based rules in real-time.
Victor Lund of WAV Group highlights a more immediate threat: agents inadvertently compromising MLS data through their daily AI tool usage. Installing a chatbot's browser extension allows the AI to view every page an agent accesses, including confidential MLS listings, without stolen credentials or system breaches. This exposure is 'invisible' because it mimics normal behavior. Copying and pasting MLS data into free chatbot accounts is another risk, as free versions often use user inputs to train large language models, unlike paid subscriptions that offer opt-out options. The financial and legal ramifications of unauthorized AI scraping are substantial; for instance, federal statutory damages for copyrighted listing photos start at $750 per violation. Carter also worries about the accuracy of AI-generated real estate advice, which could lead to costly decisions for consumers due to AI's inferential nature. California’s Department of Real Estate holds brokerages liable for advice given by their AI tools, treating them as unlicensed assistants under their control, underscoring the legal risks for brokerages and agents.
Existing MLS policies discouraging the upload of data to free AI services are largely unenforceable. Therefore, technical solutions are crucial. Bright's upcoming MCP rollout, alongside initiatives from CRMLS and vendors like FBS and Cotality, aim to plug data loopholes that facilitate AI scraping and mass downloads. Some MLSs are also incorporating explicit AI-training prohibitions into their data licensing agreements. However, Lund cautions that bans alone are ineffective due to enforcement difficulties. He advocates for technical fixes, such as preventing AI crawlers from indexing web pages and embedding immutable ownership metadata directly into listing photos, like Cotality's Trestle Defender. Sajja envisions MLSs evolving beyond data repositories to become 'trusted intelligence layers in the AI era,' allowing for innovation through compliant AI training on MLS data. Carter stresses the urgency of action, warning that continued inaction on MLS data safety in the face of rapidly advancing AI will have severe consequences for the real estate industry.