An earlier version of this research is available as an SSRN preprint.
Research question
Early-stage planning often requires consequential site decisions before complete datasets, formal indicators, and stable objectives are available. GIS-based multi-criteria decision analysis works best once indicators and weights can be specified, while conventional machine-learning and computer-vision models usually require labeled training data. Neither condition is easy to satisfy at the beginning of a coastal planning project.
This study asks whether a multimodal large language model can provide a fast, comparable, and reviewable reference under these constraints. It examines three questions: how consistent planners are with one another, how closely MLLM ratings align with expert consensus, and what happens to group consistency when AI is included as an additional evaluator.
Study design
The study sampled 357 locations along the coastline of Shanwei, a region containing sandy beaches, rocky shores, ports, and semi-enclosed bays. Each sample used a 1.5 × 1.5 km satellite image together with limited background information on tidal conditions and prevailing seasonal wind directions. Socioeconomic, policy, and other statistical data were withheld to reproduce an information-limited planning context.
Six professional planners with three to eight years of experience assessed the samples independently through a controlled web application. Each location received at least two expert ratings, producing 1,082 human judgments. The MLLM evaluated the same sites in a zero-shot setting through an automated API pipeline.
Both expert and AI assessments used a five-point scale for one holistic suitability judgment and five analytical dimensions: shoreline morphology, bay form and exposure, climatic use stability, landward support capacity, and spatial continuity and organization. The overall rating was recorded as an independent professional judgment rather than calculated as a weighted sum of the five dimensions.
The evaluation has three layers. First, each planner’s ratings are compared with the mean of the other experts to establish a human benchmark. Second, AI ratings are compared with the expert consensus using Spearman’s rank correlation and mean absolute error. Third, the MLLM is added to the group as another evaluator to test whether it changes group agreement. All three tests are repeated for the full sample and for the most and least suitable 20% of sites.

The same judgment framework is applied to expert and AI ratings, followed by three layers of validation.
Results
Across the 357 sites, the MLLM and expert panel produced similar broad spatial patterns. Both distinguished clearly suitable coastlines from heavily constrained ones, while the more ambiguous middle of the sample generated greater variation in judgment.

MLLM-generated overall suitability ratings across the sampled coastline.

Expert-average ratings provide the corresponding human reference using the same suitability legend.
The five analytical dimensions reveal why places receive different overall ratings. The model is most convincing when it can read directly visible spatial conditions, including shoreline form and adjacent land support. It is less reliable when a judgment depends on dynamic coastal processes such as exposure, tides, or seasonal wind conditions.
Validation in brief
Experts agreed most strongly on obviously suitable and unsuitable places; disagreement was concentrated in borderline cases. The MLLM followed the same broad pattern and performed within the range of variation already present among professional planners.

Expert judgments become more consistent when spatial conditions are unambiguous.
Adding the MLLM as another evaluator did not reduce overall group consistency. Its contribution was selective rather than universal, supporting a role as an additional planning reference instead of a substitute for expert judgment.

The effect of adding AI varies by dimension but remains modest overall. Detailed coefficients, uncertainty estimates, and statistical tables are available in the preprint.
Example MLLM outputs
The system returns more than a score. For each site it provides a concise overall judgment and a dimension-by-dimension explanation tied to visible coastal conditions. The two examples below show how the same framework differentiates a sheltered recreational bay from a heavily engineered industrial shoreline.

Sample 0156: the model identifies a sheltered bay, accessible shoreline, and supporting land uses as a comparatively strong candidate.

Sample 0012: the model flags hardened edges, exposure, industrial land use, and weak public access as major constraints.
Planning use
The results support a selective role for MLLMs in the initial screening of a large candidate pool. Clearly strong and weak sites can be identified for preliminary triage, while ambiguous cases are routed to expert review. Ratings for directly visible conditions, including shoreline form and adjacent land support, are more suitable for AI-assisted screening than judgments about wave exposure or seasonal climatic dynamics.
An AI-expert disagreement should therefore be treated as a prompt for additional data or a field visit. This places the MLLM before, rather than in place of, formal GIS-based assessment: the model provides an early reference while information is sparse, and GIS-MCDA becomes more useful as objectives, indicators, and datasets mature.
Limits and current status
The evidence comes from one coastal region, six planning experts, and static satellite images. The study compares rating outcomes rather than the reasoning processes behind them, and it does not establish an objective ground truth for suitability. Future work will need to test different coastal types, broaden the expert panel, and incorporate multi-temporal environmental information.
The research is ongoing, and the manuscript is under revision.