Observations for the next generation of polar prediction
Observational campaigns have always been central to improving weather prediction. The campaigns fulfil several purposes; evolving the scientific knowledge of the climate system, reveal why forecasts fail, capture processes that sustained observing networks miss and give guidance to future observational systems. As forecasting systems are becoming more integrated across the Earth system, running at higher resolution and increasingly drawing on AI technologies, future campaigns need to adapt with them. Two ECMWF events explored these questions in Reading between 29 June and 3 July 2026: the 2026 AR Recon Workshop and the 2nd Observational Campaigns Workshop for Better Weather Forecasts.
Group photo from the 2nd Observational Campaigns Workshop for Better Weather Forecasts. Photo credit: ECMWF
The first event centred on atmospheric-river reconnaissance over the North Pacific and North America; the second included seven polar talks and contributions from several PCAPS-endorsed activities. Linus Magnusson, PCAPS Steering Group’s liaison to the WMO Scientific Steering Committee, and PCAPS Steering Group member Gunilla Svensson attended in Reading, while PCAPS Steering Group member Clare Eayrs joined online to introduce the ORCAS Task Team. The presentations and recordings are available through the public workshop programme.
Forecasting is changing, and so are observational campaigns
Most AI weather models learn from reanalyses (gridded reconstructions produced using a conventional forecast model and data assimilation). ECMWF is also testing GraphDOP, which learns from observation records and predicts future observations without a separate numerical model or data-assimilation step.
Because GraphGOP skips that step, it never builds a full model state of the atmosphere. Every observation it sees shapes its forecasts directly. The system is experimental but already skilful. In tests for June 2022, its upper-air forecasts performed similarly to ECMWF’s Integrated Forecasting System during the first 24–36 hours, although the conventional system retained an advantage at longer ranges.
This approach changes the role of observations. In addition to supporting process studies, initialisation and verification, observations may provide training data and independent tests of AI systems. Direct use of observations increases the importance of calibration, uncertainty information, and metadata. If an instrument is replaced or its data reprocessed without a clear record, the AI system may learn the resulting shift in the data as real atmospheric signal, rather than as an artefact of measurement.
Campaigns planned now must therefore deliver two things: data documented well enough to stay usable as model architectures change, and independent process measurements that reveal why forecasts succeed or fail.
Campaigns reveal the process errors routine observations still miss
Standard forecast scores show where a system performs poorly but do not necessarily explain why. The polar talks showed how field campaigns can connect forecast errors to specific processes, and how forecasts can support decisions in the field.
Michael Tjernström described how the ARTofMELT expedition used three-to-five-day forecasts to decide whether to move the icebreaker Oden and establish an ice camp before an atmospheric river arrived. A virtual forecast office converted ECMWF ensemble output into daily outlooks, and the team rehearsed the entire routine ahead of the expedition.
The forecasts were not uniformly successful: one forecast placed melt onset too early. However, successive forecasts consistently captured the atmospheric-river event associated with the observed melt onset on 10 June. As Tjernström put it, “A well-prepared, rehearsed, and organised forecasting and outlook routine is really crucial for a field campaign.” The lesson is that campaign forecasting must be co-designed with operational centres before deployment.
Susanne Crewell's overview of the (AC)³ programme illustrated what sustained investment can achieve. More than 100 research flights from Svalbard and Kiruna have contributed to almost 3,000 datasets that support satellite and reanalysis evaluation and studies of Arctic clouds, cold-air outbreaks and warm-air intrusions. For example, comparing airborne measurements with satellite snowfall products, and accounting for the blind zone near the surface where satellites cannot detect precipitation, the team found that satellites underestimated snowfall by roughly a factor of two on average across five campaigns. This pattern was only visible because the aircraft data existed to check the satellite record against.
Luise Schulte used observations from MOSAiC and the Canada–Sweden Arctic Ocean Expedition 2025 to diagnose errors in Arctic cloud forecasts. ECMWF forecasts produced too few clouds containing liquid water in winter, too much liquid water in summer and too many very low clouds. In some stable conditions, the short-range forecast differed so much from balloon measurements that the quality control system flagged the good observations as suspect and gave them little weight in the analysis. The forecast’s own error, in other words, blocked the correction the observation could have provided.
Gunilla Svensson examined what happens when warm, moist air moves from open ocean across sea ice, using observations from HALO-(AC)³ and model experiments. Across ten cases, vertical ascent was the dominant contribution to cooling the arriving air and forming clouds. Cloud processes and mixing near the surface then determined how much energy reached the sea ice. Properly capturing this transformation would require tracking the same air mass for roughly 48 hours, potentially over multiple research flights. Here, the model analysis itself defined what the next campaign needs to measure.
Sarah Keeley demonstrated the complementary value of repeated, distributed observations from the SvalMIZ campaigns. Their wave observations helped ECMWF introduce waves propagating into sea ice in a new version of its forecast system. The resulting change improved forecast scores, particularly over the Southern Ocean. Early access to the observations, rigorous quality control and close cooperation between observers and model developers were central to turning campaign data into a model improvement. This work connects directly with the Distributed Observational Networks to Advance Coupled Forecasting Systems Task Team, which is examining how distributed measurements can support atmosphere–ice–ocean–wave prediction.
Together, these examples show what campaigns add. They can identify why forecasts fail, define what should be measured next and, in some cases, support operational improvements.
Observations only retain value if they remain usable
Collecting observations is only the first step. Much of a campaign dataset's value is lost when variables are difficult to find, metadata are incomplete or observations cannot readily be compared with model output.
Roberta Pirazzini presented Merged Observatory Data Files (MODFs), developed through YOPPsiteMIP. They place measurements from different instruments into a common format, along with information about how the observations were collected and processed. Matching model files use the same structure, making it easier to test how well models represent individual processes. The PCAPS Processes Task Team is extending this approach, applying merged data formats and tools across MOSAiC, Greenland Summit, Antarctic sites and icebreaker expeditions.
Standards must remain practical. A focused set of well-documented variables, uncertainties and metadata, with room for project-specific additions, is more useful than an exhaustive standard that few campaigns can implement.
Long-term stewardship must also be connected to operational use. Radiosondes and other conventional campaign observations often fail to reach forecast centres because identifiers, metadata or institutional pathways were not arranged before deployment. The transition to the WMO Information System 2.0 provides a route for sharing these observations in real time.
Delivery alone is not enough. Campaign teams and forecast centres must also confirm that observations were assimilated and given appropriate weight in the analysis. PCAPS can help by setting clear data-sharing expectations for endorsed projects, supporting practical common products and connecting campaign teams with operational centres during planning.
Linus Magnusson describing the role of WMO/WWRP, PCAPS and endorsed projects. Photo credit: ECMWF
What should campaigns measure for AI?
ORCAS connects observational scientists and AI developers around three questions: how should AI-based sea-ice forecasts be evaluated beyond standard skill metrics, how can their physical credibility be tested, and what should future observing systems measure?
One challenge is independence. Historical observations may already have influenced an AI system through its training data or model development. Future campaigns may therefore need to set aside observations specifically for evaluation, for example, keeping a set of moorings of flights out of any training dataset before a season begins, and document their relationship to training datasets from the outset.
Evaluation should also use variables beyond those routinely predicted by AI models, such as sea-ice concentration, thickness, and drift. Measurements of snow, melt ponds, ice deformation, brine content and turbulent fluxes can help test whether relationships within a forecast are physically plausible, even when those quantities are not forecast directly. Targeted campaign measurements may therefore add more value by diagnosing model behaviour than by adding volume to a training dataset.
There is unlikely to be one measurement strategy for every AI architecture. Future campaigns should provide adaptable physical benchmarks while combining intensive process studies with broader spatial coverage.
ORCAS is starting with sea-ice thickness as a pilot for developing practical recommendations. The PCAPS Sea Ice Thickness Task Team, is tackling the same problem, focusing on observational biases, uncertainties and data assimilation. The Verification Task Team is developing assessments across atmospheric, sea-ice and process-based prediction. Together, these activities can help determine what to measure and how to judge whether new systems are reliable.
Looking ahead to Antarctica InSync and IPY5
Antarctica InSync (2027–2030) and the Fifth International Polar Year (2032–2033) are the next major international planning windows for polar observations. Planning must start now. Ship and aircraft campaigns need years of logistical preparation, while autonomous platforms, satellite underflights and ground-based networks must be coordinated across programmes and nations. Data formats, real-time exchange and independent evaluation strategies cannot be decided once instruments are already in the field.
AR Recon offers a useful organisational model. Its repeated aircraft deployments are supported by sustained coordination among operational centres, universities, campaign teams and funding agencies. Scientific priorities and observing infrastructure have developed together.
The immediate priorities are:
Agree which physical processes Antarctica InSync and IPY5 should test.
Set aside independent observations for evaluation before campaigns begin.
Define a practical core of shared variables, uncertainty information and metadata.
Link campaigns to operational centres and real-time data pathways wherever possible.
Coordinate ship, aircraft, satellite, autonomous and ground-based observations across national programmes.
Make long-term data stewardship part of campaign design and funding.
New models alone will not improve polar forecasts. Progress also depends on observations that reveal why systems fail, test them in unfamiliar conditions and remain usable after campaigns end. PCAPS can help turn lessons from individual projects into shared observing priorities for Antarctica InSync and IPY5, and connect those observations to prediction systems and services.

