Can AI Help Regulators and Market Participants More Effectively Analyze Thousands of Rule Proposal Comments?

Posted Monday, September 21

By Campbell Pryde, President and CEO, XBRL US

Summary: 

  • Agencies and market participants spend significant time actively reviewing a high volume of rulemaking comment letters. 
  • Current agency comment letters are unstructured and require manual reviews. The increasing use of AI tools custom designed to review and extract information from specific comment types increases opacity and cost for regulators and other stakeholders.
  • Structured templates to express comment requests and responses would enhance transparency, reduce cost, and improve analysis by AI tools used by regulatory staff and market participants.

U.S. regulatory agencies publish 3,000 to 4,500 final rules each year, according to the Office of the Federal Register. Each rule is typically preceded by a Notice of Proposed Rulemaking (NOPR) — a draft rule posted during a public comment period to collect market input. In 2018 alone, regulators published 2,098 NOPRs and finalized 3,368 rules. Under the Administrative Procedure Act (APA), which governs federal rulemaking, agencies must publish rule proposals in the Federal Register and weigh the recommendations submitted by market stakeholders before issuing a final rule.

Rule proposals are often lengthy and dense, historically ranging from 50 to well over 200 pages, and each can draw hundreds or even thousands of comment letters that regulatory staff and market participants sift through and analyze. Those letters, too, are typically long, with detailed responses to questions posed, and are submitted as unstructured PDF documents. One recent example: an SEC proposal on semiannual reporting drew more than 181,000 unstructured comment letters (count based on the analysis at the time of this writing), according to an AI-powered tracker built by Professor Tzachi Zach of The Ohio State University's Fisher College of Business.

Regulators and market participants need consistently structured input to enable AI tools to more effectively analyze NOPR responses more efficiently — structured machine readable templates and artificial intelligence could play an important role.

Why Is Structured Data Better Than Unstructured Data for AI Applications?

Over the past 15 years, the SEC has steadily built a pool of highly contextualized, structured financial data — both quantitative and qualitative — reported by public companies and investment management companies. That pool of semantically structured data is now a rich resource for powering efficient, reliable, and cost-effective AI tools. Commercial data aggregators and analytics providers already extract this data with ease, serving up results through traditional analytics and increasingly, using LLM tools that query and extract structured data directly. 

AI performs significantly better when working with structured data versus unstructured data. It finds consistent semantic meaning in structured XBRL data by reading the highly contextualized “wrapper” around each value which carries the name, label, definition, units of measure and other dimensional data embedded in the value itself. With PDF files, AI derives meaning by parsing the value from a table and attempting to read columns and rows, or from narrative text using OCR (optical character recognition) or character extraction and then inferring structure. Unstructured content is simply not efficiently captured by AI because of the work involved to discern meaning and the ambiguity often present. That added work and effort can also produce inaccurate results.

Efficiency matters more every day, as the energy and infrastructure costs of running LLMs continue to climb. AI firms have been subsidizing usage heavily to encourage adoption, but the “AI free lunch” won’t last: as investors look for returns, the unlimited private equity and venture funding that has absorbed AI’s high deployment costs will dry up. 

And while market competition from new LLM entrants is likely to put downward pressure on costs, that means less consistency in output. When LLMs work with unstructured data, the same prompt made to various LLMs, from Claude to ChatGPT to Gemini, will most likely produce different outputs, giving the market an inconsistent understanding of financial market economics. The more LLMs in use, the more inconsistencies will arise. Conversely, when LLMs work with structured data anchored to a digitally consistent semantic data model, the richer context provided to the LLM makes the output generated consistent across AI platforms. 

Regulators that already maintain structured XBRL data programs — the SEC, the Federal Deposit Insurance Corporation (FDIC), and the Federal Energy Regulatory Commission (FERC) — have a big head start on other U.S. agencies in using AI, simply because of the structured data they’ve been collecting for years.

How Can AI Powered by Structured Data Improve Comment Letter Analysis?

Fast, efficient processing of comment letters is critical to how regulators do their work. Evaluating comment letters should be a natural use case for AI, but most comment letters provided to U.S. agencies arrive as unstructured PDFs that LLMs cannot easily or accurately parse. 

Some non-U.S. regulators already require comment letters to be submitted through structured online forms, making submissions far easier to inventory and process not only for the regulatory staff but also for all market participants. The European Securities and Markets Authority (ESMA), for example, provides a Microsoft Word template for commenters to fill in. When responses follow a structured questionnaire, the resulting data can be processed as structured, meaningful information — without the need to parse as free text.

Can XBRL Optimize the Reliability and Efficiency of Comment Letter Analysis?

U.S. regulators could take this a step further with a structured questionnaire that produces data natively in XBRL. Each field could be tied to a clearly defined XBRL “tag” carrying an unambiguous definition, an entity identifier such as the Legal Entity Identifier (LEI), a label, and dimensional information — for example, commenter type (law firm, software company, issuer, etc.). With comment letter data in structured format, following a consistent semantic data model, regulators could query even 181,000 letters at once, more confidently, more timely and at far lower cost.

Publishing that data in structured format would also give other users — law firms, journalists, academic researchers, investors, and issuers — richly detailed information to predict rulemaking outcomes and gauge market sentiment at a granular level. And because XBRL is a widely adopted standard, thousands of commercial and open-source tools already exist to extract and process XBRL-formatted data, keeping processing easy and costs low. 

The difference between using a widely adopted standard like XBRL versus building a custom XML or JSON schema, is that a custom schema requires the creation of custom-designed reporting, collecting, extraction, and analytical tools, exponentially increasing the cost of implementation. Custom schemas produce data in a custom structure so data cannot be combined, shared, or inventoried with other data. Custom schemas do not offer the economies of scale (and lower costs) or data interoperability that following an open, widely used semantic data model like XBRL provides.

Can the SEC Reduce the Cost and Increase the Efficiency of Its AI Use?

Many U.S. regulators are already turning to LLM tools. The SEC publishes an AI Use Case report on its “AI at the SEC” webpage; a review of that report shows 49 distinct use cases spread across 22 divisions, relying on a wide range of LLMs and commercial vendors with different AI solutions targeting enhancement of comment letter review. The table below consolidates those use cases and gives examples of the tools in use.

That variation — in use cases, tools, and platforms — suggests each division often builds its own AI applications, and may be using different tools for the same purpose. Comment letter analysis is a good example: it’s a task many SEC departments perform independently, and the Commission would gain efficiency, reliability, and cost savings if all of them used the same approach and toolset.

Collecting all SEC comment letter data in a structured, machine-readable format — under one consistent semantic data model, paired with a Commission-wide AI toolset — would lower the cost of preparing regulations and analysis by market participants while surfacing richer, more granular information to inform better final rules.

Comment letters hold a wealth of information — not just for regulators, but for the broader market seeking to understand sentiment and trends around proposed changes to disclosure and market structure. Structured data following a consistent semantic data model like XBRL, coupled with AI, is a win for regulators and market participants alike.

Comment