Nobody Uses Your Open Data Portal. Search Is Why.
A journalist covering the city council wants to know how the police department's use of overtime has changed over the past three fiscal years. Your open data portal contains this information. It is split across four datasets with names like 'Personnel Expenditure Detail FY23-Q4' and 'Sworn Officer Time Reporting'. The journalist searches 'police overtime'. The portal returns 78 results, ranked by upload date. She spends 20 minutes clicking through, then gives up and files a FOIA request.
This scene plays out every day at every municipal open data portal in the country. It is not a data problem. It is a findability problem. And it is undermining exactly the accountability, transparency, and civic engagement outcomes that the open data program was created to deliver.
Why Portal Search Was Never Good Enough
Most municipal open data portals run on one of a handful of platforms: Socrata, CKAN, ArcGIS Hub, or a custom build. All of them ship with search interfaces designed by people who already know how open data catalogues work. The search assumes the user knows the metadata schema, understands the difference between a dataset and a view, and can navigate faceted filters.
The average resident visiting your open data portal has none of that context. They arrive because they saw a news story that mentioned an open dataset, or because a city council candidate cited your portal in a speech, or because a civic tech volunteer sent them a link. They type a question in ordinary English. The search returns nothing useful. They bounce.
Our portal has more than 400 datasets. We know from analytics that the median session visits fewer than three of them before leaving. The datasets people are looking for are almost always there. The search just cannot map from what people type to what the datasets are called.
Open data program manager, mid-size city government, West Coast
The Naming Gap
The heart of the problem is a mismatch between how datasets are named and how residents describe what they want. Datasets are named according to the source system, department, and internal conventions of the agency that publishes them. Residents describe what they want using the language of the outcome they care about.
Consider a resident searching 'is my street getting resurfaced this year'. The relevant dataset might be called 'FY26 Capital Improvement Program - Public Works Division 4 Pavement Management Plan'. There is no keyword overlap between the query and the dataset name. Traditional search returns nothing. The resident has no way to know that dataset exists.
This is exactly the problem AI Search solves. An AI search layer over the portal catalogue understands that 'street resurfacing' semantically matches 'pavement management', that 'this year' matches 'FY26', and that 'my street' means the resident wants to filter by geography. It returns the dataset with a synthesised summary explaining what it contains, and offers a follow-up path to see just the records for the resident's neighbourhood.
Three Layers of Findability the Portal Actually Needs
Layer 1: Dataset Discovery
The first layer is finding the right dataset when a user does not know its name. This is where AI search over the catalogue metadata delivers the fastest improvement. The AI indexes every dataset's title, description, tags, columns, and update history. When a user searches, the AI matches semantic intent to the datasets that could answer their question, then returns those datasets with a plain-language explanation of what each contains.
Layer 2: Row-Level Search
The second layer is finding specific records inside a dataset. A user who has landed on the right dataset then wants to know 'what does this dataset say about my address' or 'show me the records for the police department for last quarter'. Traditional portal filtering supports this if the user understands the columns. AI search supports it in natural language: the AI translates the question into the appropriate column filters and returns the answer.
Layer 3: Cross-Dataset Question Answering
The third layer is the hardest and the most valuable. A journalist asking 'how has police overtime changed over the past three years' needs an answer that spans multiple datasets. AI search that can traverse the catalogue, combine the relevant records, and produce a synthesised answer with citations to the underlying datasets is genuinely new capability for open data portals. This is where civic transparency stops being a data dump and starts being a conversation.
400+
typical dataset count on a mid-sized city open data portal
62%
of portal search sessions end without a dataset download
3
median datasets visited per session before the user leaves
18 mo
average age of a dataset name that no longer matches how residents describe it
5x
increase in dataset downloads observed after portal AI search deployment in early pilots
70%
of FOIA requests in one large city were for information already available on the open data portal
What This Costs the City When It Fails
The most visible cost is the FOIA request volume. A resident, journalist, or advocacy organisation that cannot find the information they need on the portal often files a formal records request instead. Each FOIA request costs the city between $300 and $1,500 to process, depending on complexity. A portal that surfaced the underlying dataset would have handled the same request at essentially zero marginal cost.
The less visible cost is the erosion of the transparency mandate itself. Open data programs were funded in most cities to advance civic engagement, government accountability, and third-party innovation. When the portal is unusable in practice, the program stops delivering those outcomes. Political support for the program erodes. Budgets get cut. The datasets stay online but stop being maintained.
The Deployment Path That Actually Works
Most cities that have improved open data portal findability have done it in three phases.
- 1Phase one: metadata cleanup. Audit dataset names, descriptions, and tags. Rewrite titles that are internally jargon-heavy. Add plain-language summaries. This work is not glamorous, but it improves both traditional search and AI search at once. Budget 15 to 30 minutes per dataset.
- 2Phase two: AI search over the catalogue. Deploy AI search that indexes the metadata and returns natural-language results. Prioritise the top 100 most-viewed datasets. This is the visible upgrade residents notice.
- 3Phase three: cross-dataset question answering. Once the AI is performing well on catalogue queries, extend it to synthesised answers that span multiple datasets. This is the accountability journalism layer. It usually requires a per-dataset configuration to define which columns are safe to aggregate and how citations should be presented.
Data Quality Guardrails
Open data brings its own accuracy stakes. If AI search on the portal reports that a dataset shows X, and the dataset actually shows Y, the reputational damage is significant. Deploy AI search on open data with three guardrails in place.
- Citation to the underlying dataset and, where possible, to the specific rows the answer is drawn from. Users should be able to click through and verify the source.
- Timestamps on the source data. A dataset that was last updated three years ago should not be presented as current. The AI should surface the update date alongside any answer.
- Confidence signalling. When the AI is stitching together multiple datasets or making an inference, the response should say so plainly rather than presenting the answer with the same confidence as a direct lookup.
Where This Sits in the Broader FOIA and Transparency Program
Portal findability is one part of a larger transparency story. The other pieces are FOIA processing efficiency (covered in our FOIA processing blog post) and general public records search on the main city website. All three pieces work together. A resident who can find data on the portal does not file a FOIA. A resident whose FOIA is answered promptly does not need to escalate. A journalist who can search across public records and datasets in a single interface produces better reporting, which the city ultimately benefits from.
The one dataset test
Open your portal in an incognito window. Search for 'police overtime', or 'street resurfacing', or 'business license by neighbourhood', whichever is most politically relevant in your city right now. Time how long it takes to reach the relevant dataset. If it takes more than 30 seconds, that is your baseline. AI search should cut it under 10.
See the full digital experience picture
Download the Government Digital Experience 2025 Report for benchmark data on how residents actually navigate state and local government sites, portals included.
Related reading
Pull your portal analytics for last quarter, then book a demo and we'll show you what AI search would surface for your highest-traffic queries.
Ready to see it in action?
Book a demo and we'll configure Keyspider on a live sample of your content, within 48 hours.
Book a Demo