ICU Management & Practice, Volume 26 – Issue 3, 2026

img PRINT OPTIMISED
img SCREEN OPTIMISED

This article traces evolution of critical care data from early ICU registries use analogies from other industries to modern open databases and future data fabrics. The article highlights Mayo Clinic ICU Data Mart as a clinically driven project of near-real-time data reuse for surveillance, decision support, reporting and research. Looking towards 2050, the authors suggest that ICU data must become multimodal real-time, flexible to be federated but semantically standardised. Mac3D-like platform is presented as a bridge toward a learning ICU environment that supports bedside expertise and compassionate care.

 

Electronic health records clearly improved access to clinical information, but they have not solved deeper data problems in intensive care. The Intensive Care Unit (ICU) is one of the most data-rich environments in medicine: monitors, ventilators, infusion pumps, laboratory systems, pharmacy platforms, radiology, blood banks, nursing flowsheets, clinical notes, microbiology systems, and transport records all generate clinically meaningful information (Pickering et al. 2013). Recent EHR adoption has been a step forward in organising that data into usable user interfaces; however, “data silos” still exist across different vendor systems. Data are often delayed by batch processing or stored in formats that are difficult to query, harmonise, or reuse. The central challenge of ICU informatics is therefore not simply data generation. The gap between “data generated” and “data usable” is the central story of ICU informatics.

 

As we have seen during the 1990s, many industries began building domain-specific data repositories (Datamarts) long before the term “Big Data” became common. Telecommunications companies mined call-detail records for fraud detection and customer profiling, retailers used scanner data and loyalty cards to understand purchasing behaviour, FedEx and UPS transformed parcel movement into trackable digital events, airlines used reservation databases to manage fares and availability, financial markets created high-frequency trade-and-quote archives, and geospatial and weather agencies built massive sensor and image repositories. These systems were not generic IT projects. They were industry-specific data infrastructures that converted routine operational transactions into prediction, surveillance, benchmarking, and decision support.

 Herasevich t1

Early Critical Care Repositories

Early database pioneers in Critical Care included the Intensive Care National Audit & Research Centre (ICNARC), established in the UK in 1994, and the Australian and New Zealand Intensive Care Society (ANZICS) CORE, established in 1992. These organisations were created by clinicians to standardise data collection and elevate patient care. Focusing on standardised benchmarking and risk-adjusted mortality modelling, they contributed to the development of early scoring systems to predict patient outcomes based on the severity of illness and physiology. Other use cases included standardisation and clinical audit. What started as simple tracking systems for mortality and length of stay has evolved into the backbone of critical care research by providing infrastructure for multicentre clinical trials.

 

In the USA, Project IMPACT (Improved Methods of Patient Information Access of Core Clinical Tasks) was driven by the Society of Critical Care Medicine. In 1996, a database was created to collect standardised critical care data across participating ICUs. Later, it was used to create the Mortality Probability Model (MPM) benchmarking system. By the early 2000s, Project IMPACT had become a major source for USA critical care outcomes research supporting studies of mortality prediction, ICU readmissions, staffing models, resource use, invasive monitoring and variation in practice.

 

The importance of these historical databases is that they demonstrate how critical care used standardised multicentre data for risk-adjusted comparison long before modern EHR-derived data marts and open ICU databases became common. However, those projects were registry-like rather than real-time. They relied on trained abstraction, captured selected variables and did not include the high-resolution device, waveform, alarm and workflow data that future ICU data fabrics will require.

 

The Mayo ICU Data Mart

The ICU Datamart movement in the 2000s represented a delayed entry into the same data-infrastructure revolution that transformed retail, logistics, finance, telecommunications and travel in the 1990s.

 

The ICU Datamart at Mayo Clinic, Rochester, Minnesota, USA, emerged in 2006 as a pragmatic response to this challenge. Built "one brick at a time", by 2011 Mayo’s ICU Data Mart contained near-real-time data for 206 ICU beds, approximately 15,000 ICU admissions per year and 200,000 vital records per day. It included historical data going back to 2003 and, most importantly, accumulated data with only a 15-60 minute delay from real time. This created opportunities to develop and test real-time syndrome surveillance and clinical decision support systems. The ICU Datamart at Mayo Clinic was also used for clinical research data extraction, practice reporting and traditional modelling of critical care illnesses (Herasevich et al. 2010).

 

The 2011 article in Healthcare Informatics (Herasevich et al. 2011) is historically important because it framed several principles that remain relevant today. The first was the “Lego” principle - build usable database components incrementally rather than waiting years for a perfect end solution. The second was the “UNIX” principle - avoid unnecessary gatekeeper user interfaces and allow skilled users to query data directly through statistical and analytic tools. The third was the “Matrix” principle - preserve meaningful raw data whenever possible because premature preprocessing can destroy future analytic value.

 

ICU Datamart continues today on a cloud technology platform for scalability, reliability, and high-performance, with a one-day delay from real time and 23 years of accumulated ICU data. It contains data from more than 160,000 ICU patients and over 350,000 ICU admissions. This makes it one of the largest comprehensive clinical data repositories in the world.

  

Public ICU Databases and the Reproducibility Era: MIMIC and AmsterdamUMCdb

Another major step was the creation of de-identified critical care databases. These datasets changed the field by making ICU data available to investigators beyond a single institution. They enabled reproducible research, external validation, benchmarking, education and international data science communities. MIMIC is the most influential example. The current MIMIC-IV release on PhysioNet is a de-identified dataset of patients admitted to the ICU at Beth Israel Deaconess Medical Center, Boston, Massachusetts, USA. MIMIC-IV contains data from more than 90,000 ICU patients. Version 3.1 was released in October 2024 and made available on BigQuery, reflecting the movement from downloadable research files toward cloud-based analytic environments (Johnson et al. 2023).

 

AmsterdamUMCdb added an important European perspective. It is the first freely accessible European intensive care database and contains de-identified data from more than 23,000 ICU and high-dependency unit patients admitted between 2003 and 2016 to Amsterdam University Medical Center, Netherlands. It consists of 7 tables and includes clinical observations, vital signs, scoring systems, device data, laboratory results and medications. The current AmsterdamUMCdb release uses OMOP Common Data Model version 5.4. The original AmsterdamUMCdb publication also emphasised responsible data sharing, de-identification, contractual governance and privacy review as necessary infrastructure for international critical care data science (Thoral et al. 2021).

 

Those two international open ICU databases created the reproducibility era. MIMIC and AmsterdamUMCdb made de-identified critical care data available to credentialed researchers for machine-learning research and education. These databases also reveal the limitations of the Datamart paradigm. They are primarily retrospective, static and incomplete representations of the ICU environment. Although they are invaluable for discovery, they were not designed to function as the real-time operational system of the ICU of the future.

 

Modern Bottlenecks Enabling Future Needs

EHR-based data marts have transformed ICU research and reporting. They now represent an important bottleneck for the next era of critical care analytics. Their greatest limitation is that they remain largely EHR-centred. They capture what is documented, ordered, billed and reported, but not necessarily what is happening continuously at the bedside. High-frequency waveforms, ventilator curves, infusion pump events, alarms, RTLS signals, device metadata, computer vision and acoustic output, biosensors and genomic data often remain outside the traditional EHR or are trapped in separate vendor systems.

 

In many institutions, data access is still delayed by batch processing. Semantic mapping is often incomplete, and local terminologies make cross-system or multicentre analysis difficult.

 

The result is a paradox. Modern ICUs generate enormous quantities of data, but EHR-derived data marts provide a delayed, low-resolution and locally idiosyncratic representation of critical illness.

 

From Datamart to Data Fabric

The ICU of 2050 most likely will represent a paradigm shift from a reactive, fragmented clinical setting to a proactive hospital-wide environment. Future ICU management will require the seamless convergence of the Internet of Medical Things (IoMT), immersive human-computer interfaces, human-in-the-loop clinical decision support and closed-loop autonomous artificial intelligence.

 

By 2050, the ICU data mart will likely no longer be a "mart" in the traditional sense. It will be a living data fabric – a federated, continuously updated, semantically standardised environment that connects scattered data sources such as the bedside EHR, devices, registries and research databases, models and clinical workflow data.

 

The future ICU data platform will have several defining features. First, it will be multimodal. It will integrate not only structured EHR data but also waveforms, ventilator data, infusion pump data, alarms, imaging, ultrasound clips, procedure videos, genomics, microbiome data, bedside biosensors, wearable data, RTLS location data, acoustic signals, computer vision data and many other data streams from the care environment. These data will not be collected simply, even though they are technically available. They will be curated to explain physiology, workload, treatment response, safety and recovery.

 

Second, data will be real-time or near real-time. Retrospective data will remain essential for initial research, but true clinical decision support requires real-time data.

 

Third, the future ICU data platform will be semantically standardised. A potassium result, vasopressor dose, ventilator mode, sepsis diagnosis, ECMO run or delirium assessment must mean the same thing across units, hospitals, countries and time. This is why common data models, terminologies and critical care data dictionaries matter. Currently, there are many different standards in use. The reality is that there is no, and probably never will be, single universal standard for healthcare data. SNOMED CT, ICD, HL7 FHIR, OMOP and others have unique features and different purposes. For example, the Society of Critical Care Medicine published its own core Critical Care Data Dictionary with common data elements across nine domains, providing an important foundation for harmonised critical care data collection, but it does not encompass multimodal, real-time data generated outside of the EHR (Murphy et al. 2025). New approaches to metadata, management, missingness indicators, version control and model monitoring will need to be incorporated as core functions into the Datamart of the future.

 

The next generation of ICU Datamarts must therefore move beyond EHR data extraction. The Mayo Clinic Critical Care Datamart (Mac3D) represents one example of this transition. It is a federated platform with centralised storage of critical elements designed to integrate EHR data, outputs from multiple devices, waveforms and other non-EHR metadata, while supporting on-the-fly registry mappings and preparation for future AI, machine learning and large language model, and agentic applications.

 

Mac3D-Like Systems as a Bridge to 2050

Mac3D-like systems represent the bridge between today’s EHR-derived datamarts and the ICU data fabrics of 2050.

 

Mac3D is a real-time centralised data repository supported by a federation infrastructure. The platform is currently under active development for the creation and testing of advanced clinical decision support, artificial intelligence, machine learning and large language model tools. Mac3D is also designed to support outcomes research and quality improvement.

 

The second distinguishing feature of Mac3D is the inclusion of high-resolution non-EHR data like device waveforms, biosensors, genomics and other similar data streams.

 

The third distinguishing feature is on-the-fly semantic mapping to existing standard terminologies and models such as SNOMED CT, OMOP, LOINC, and RxNorm. The system provides output in established de facto standard schemas such as MIMIC.

 

The fourth important feature is user access. Datamart itself should reside on the most advanced technical infrastructure available and support multiple modes of access. Old school SQL and ODBC/JDBC connections remain essential, but new natural-language query tools powered by large language models should provide clinicians and researchers with the ability to ask practical questions without needing to become database engineers.

 

Instead of writing complex queries, a user might ask a question in natural language, such as: “Show me all mechanically ventilated patients from the last year who received prone positioning within six hours. Include ARDS physiology, vasopressor requirements, and renal function in the output.”

 Herasevich f1

Additional Components of Future Infrastructure

Mac3D-like systems would benefit from the following additional components and capabilities.

  1. Universal IoT Gateways: Device fragmentation may be reduced through edge-computing gateways that translate proprietary device outputs into standardised, time-synchronised data streams using technologies such as OPC Unified Architecture (OPC UA). These edge gateways will help aggregate and filter thousands of sensor readings per second locally, significantly reducing cloud transmission loads while retaining high-fidelity waveform data for deep learning.
  2. Semantic Web: Future data marts may gradually abandon flat tables in favour of dynamic Semantic Web technologies (such as RDF and OWL) to create highly interconnected knowledge structures. Automated metadata intelligence will instantly map operational streaming data (like HL7 FHIR) to observational research standards (like OMOP) with high precision, linking a patient’s genomic information, real-time haemodynamic waveforms, and pharmacologic responses into an accessible knowledge framework.
  3. AI-Driven CDS: In the coming years, rule-based alerts will transition to "Context-Aware Memory" architectures powered by AI agents. These systems will provide high-speed access to a durable, continuously updated body of institutional and patient knowledge. Decision support systems will have high-speed access to a durable, continuously updating body of institutional and patient knowledge. CDS will utilise multimodal deep learning combining Convolutional Neural Networks (CNNs) to extract local temporal patterns from raw waveforms, and Transformer encoders to capture long-range clinical dependencies.
  4. Digital Twins Infrastructure: Larger datasets and advances in deep learning (DL) will allow clinicians to simulate individual patient pathophysiology in real time. These models may detect subtle physiological deterioration, predicting critical events in a near-real-time manner based on digital twin patterns.
  5. Support Federated Systems: A Mac3D-like system can serve as a node within a massive, globally connected network. Because global privacy laws most likely will continue to restrict the cross-border movement of raw, identifiable patient data, global innovation will be driven by Federated Learning. In a federated ecosystem, multinational hospitals will train AI models locally on their own encrypted datamarts. Instead of sharing patient data, hospitals will share only updated model parameters while using standardised data schemas. This type of connectivity will allow institutions to benefit from collective algorithmic intelligence and predictive accuracy trained across millions of patients worldwide.

 

The success of the ICU data fabric will not be determined by storage capacity or model performance alone. It will depend on human factors and user education. Success will depend on whether nurses trust the data, whether physicians understand model limitations, whether respiratory therapists can correct device mappings, whether alerts reduce rather than add cognitive burden and whether patients and families trust how data are used. In 2050, ICU data governance will need to be treated as a patient safety function. Data quality failures, model drift, biased predictions, broken device feeds and poorly designed alerts will not be technical inconveniences - they will be real clinical risks.

 

Conclusion

The original ICU Data Mart demonstrated that a small, clinically driven, incremental system could overcome data silos and create practical value. MIMIC and AmsterdamUMCdb showed that shared critical care data could democratise research and reproducibility. Mac3D-like platforms point toward the next stage: real-time, federated, multimodal, standards-based, AI-ready critical care data ecosystems.

 

The ICU of 2050 will be a true electronic learning environment. However, its essential function will still depend on bedside expertise, teamwork and compassion. These human qualities will be supported, not replaced, by a new kind of data infrastructure. The digital fabric of the hospital, woven from foundations established in the ICU, will remember what has happened, understand what is happening, anticipate what may happen and help clinicians evaluate what to do next.

 

Conflict of Interest

None.


References:

Herasevich V, Kor DJ, Li M, et al. ICU data mart: a non-IT approach. A team of clinicians, researchers and informatics personnel at the Mayo Clinic have taken a homegrown approach to building an ICU data mart. Healthc Inform. 2011;28(11):42, 44-5.

Herasevich V, Pickering BW, Dong Y, et al. Informatics infrastructure for syndrome surveillance, decision support, reporting, and modeling of critical illness. Mayo Clin Proc. 2010;85(3):247-54.

Johnson AEW, Bulgarelli L, Shen L, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data. 2023;10(1):1.

Murphy DJ, Anderson W, Heavner SH, et al. Development of a core critical care data dictionary with common data elements to characterise critical illness and injuries using a modified Delphi method. Crit Care Med. 2025;53(5):e1045-e1054.

Pickering BW, Gajic O, Ahmed A, et al. Data utilisation for medical decision making at the time of patient admission to ICU. Crit Care Med. 2013;41(6):1502-10.

Thoral PJ, Peppink JM, Driessen RH, et al. Sharing ICU patient data responsibly under the Society of Critical Care Medicine/European Society of Intensive Care Medicine Joint Data Science Collaboration: the Amsterdam University Medical Centers Database (AmsterdamUMCdb) example. Crit Care Med. 2021;49(6):e563-e577.