Maps and GIS · 28 Sep 2026

What is agentic GIS?

← Maps and GIS

Agentic GIS is a geographic information system operated by an AI agent. The user asks a spatial question in plain language, such as which stands within two kilometres of a road were thinned in the last five years. The agent plans the steps and chooses the GIS tools. Then it runs them and returns a map or a table. It is real, and it is early. In a strict 2026 test built from questions GIS practitioners actually asked, scored against exact outputs, the best agent fully solved about a third of the tasks. Most of the models tested produced outputs close to the correct answer.

What the usual answer says

Descriptions of agentic GIS promise spatial analysis without the specialist: ask the map a question, get the answer. A 2026 synthesis of the research, one of the top results, does list limits in spatial reasoning, execution and validation. What the descriptions rarely include is a measured rate of exactly right answers, and which answers go wrong.

GISAgentBench, 2026 preprint: 349 multi-step GIS tasks Best agent, share of tasks completed under strict tolerance-aware scoring 32.7% Most models produced outputs close to the ground truth. Geometry was recovered far more reliably than numeric attributes and complete row sets. all tasks
Under strict scoring against exact outputs, the best agent fully solved about a third of the tasks. Close is not correct: the shapes can look right while the numbers behind them are out. Diagram: Quarri.

What the benchmarks measured

GISAgentBench, a preprint from August 2026 by Abhinav Pothuri and colleagues, took "349 multi-step GIS tasks curated from GIS Stack Exchange" and ran them on real public data. Each task came with an exact correct output. Results could therefore be scored by matching outputs within a tolerance, rather than by a model's judgement. The finding: "the best agent completes only 32.7% of tasks under strict tolerance-aware scoring, although most models produce outputs that are close to the ground truth". Strict one-shot scoring is itself a choice. A specialist working with an agent over several turns would correct some of those near misses, so the figure describes today's agents on their first try, not a ceiling.

GeoBenchX, a 2025 preprint by Varvara Krechetova and Denis Kochedykov, tested eight commercial models, each given 23 geospatial functions, on multi-step tasks. Some tasks were deliberately impossible to solve. The authors report that "Common errors include misunderstanding geometrical relationships, relying on outdated knowledge, and inefficient data manipulation". One model, the authors wrote, "due its preference to provide any solution rather than reject a task, proved to be less accurate." Even the easiest group of tasks, merging and mapping data, was not reliable: "almost all models achieved accuracy over 0.6".

GeoAgentBench, an April 2026 preprint, built a sandbox of 117 GIS tools and 53 analysis tasks. The authors start from the view that "precise parameter configuration is the primary determinant of execution success". In our reading, that means details such as the buffer distance, the coordinate system and the join, which are exactly what a plain-language request leaves out.

Close is the problem

"Close to the ground truth" is a reassuring phrase with an awkward meaning for a forest business. GISAgentBench's full text explains it: "agents recover geometry far more reliably than numeric attributes and complete row sets", with "geometry and topology errors and CRS misalignment the most damaging". Vague requests "produce syntactically valid but semantically wrong outputs rather than failed calls". The shapes on the map look right while the hectares or tonnes behind them are out, and nothing in the output says so.

The fix is to ask the agent to show its work in a form a specialist can read quickly. That means the layers it used, the distances and projections it set, and the number of features at each step. A checker who sees that a buffer was set at 200 metres rather than two kilometres, or that 40 parcels went into a join and 37 came out, can catch the error quickly.

What it is useful for today

An agentic GIS is useful as a drafting tool for spatial analysis. It turns a question into a first set of steps and a first result for a specialist to check. It is not yet a replacement for the specialist, whose job shifts to checking. Are the right layers used? Is the distance right, the coordinate system consistent, the join complete? For decisions that rest on area or volume, such as a harvest plan or a sale, that check is the difference between a useful draft and an expensive error.

Quarri's own agentic GIS work is in design, not delivered, and this piece describes the field rather than any product.

When it doesn't apply

Simple spatial lookups, such as the area of one stand, carry less ambiguity, though GeoBenchX suggests even simple tasks still fail often enough to need a glance. Exploratory questions, where a rough map is enough to decide where to look, can take an agent's first answer as it is. And in an organisation without GIS expertise, there is nobody to check the agent. There the right use is narrower: questions whose answers can be checked some other way.

Quarri for forest management is built around how a forest operation runs, from the cruise to the settled account.

Sources

  1. Pothuri, Jiang, Xu and Yang, "GISAgentBench: A Practitioner-Sourced Benchmark for Evaluating LLM Agents on GIS Tasks", arXiv 2608.01645, 3 August 2026: arxiv.org
  2. Krechetova and Kochedykov, "GeoBenchX: Benchmarking LLMs in Agent Solving Multistep Geospatial Tasks", arXiv 2503.18129, 2025: arxiv.org
  3. Yu et al., "GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis", arXiv 2604.13888, 15 April 2026: arxiv.org

Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.

See it on your own data.

Live in two weeks, on the systems you already run.