Completed Computing & AI Arts, Culture & Design

VADA: Value Added Data Systems -- Principles and Architecture

In plain English

AI plain-English summary

Data scientists spend 50% to 80% of their time just cleaning and reorganising messy data before they can analyse it at all. That data wrangling bottleneck, caused by the sheer volume, speed, variety, and uncertainty of modern data, wastes enormous economic potential—estimated at £40 billion in 2017 alone. The VADA programme aims to build a new generation of data management tools that automatically handle these tasks by understanding both what the data is and what the user actually needs. Instead of forcing analysts to manually check quality and provenance, the system would track data’s origin, reliability, and cost, then deliver the best available results along with a clear explanation of their limitations. Users could feed back on results, allowing the system to continuously improve. If successful, this would radically accelerate data-driven innovation across healthcare, finance, e-commerce, smart cities, and telecommunications—any sector where analysts currently waste the majority of their time wrangling data instead of extracting insights from it.

View original technical description
Data is everywhere, generated by increasing numbers of applications, devices and users, with few or no guarantees on the format, semantics, and quality. The economic potential of data-driven innovation is enormous, estimated to reach as much as £40B in 2017, by the Centre for Economics and Business Research. To realise this potential, and to provide meaningful data analyses, data scientists must first spend a significant portion of their time (estimated as 50% to 80%) on "data wrangling" - the process of collection, reorganising, and cleaning data. This heavy toll is due to what is referred as the four V's of big data: Volume - the scale of the data, Velocity - speed of change, Variety - different forms of data, and Veracity - uncertainty of data. There is an urgent need to provide data scientists with a new generation of tools that will unlock the potential of data assets and significantly reduce the data wrangling component. As many traditional tools are no longer applicable in the 4 V's environment, a radical paradigm shift is required. The proposal aims at achieving this paradigm shift by adding value to data, by handling data management tasks in an environment that is fully aware of data and user contexts, and by closely integrating key data management tasks in a way not yet attempted, but desperately needed by many innovative companies in today's data-driven economy. The VADA research programme will define principles and solutions for Value Added Data Systems, which support users in discovering, extracting, integrating, accessing and interpreting the data of relevance to their questions. In so doing, it uses the context of the user, e.g., requirements in terms of the trade-off between completeness and correctness, and the data context, e.g., its availability, cost, provenance and quality. The user context characterises not only what data is relevant, but also the properties it must exhibit to be fit for purpose. Adding value to data then involves the best effort provision of data to users, along with comprehensive information on the quality and origin of the data provided. Users can provide feedback on the results obtained, enabling changes to all data management tasks, and thus a continuous improvement in the user experience. Establishing the principles behind Value Added Data Systems requires a revolutionary approach to data management, informed by interlinked research in data extraction, data integration, data quality, provenance, query answering, and reasoning. This will enable each of these areas to benefit from synergies with the others. Research has developed focused results within such sub-disciplines; VADA develops these specialisms in ways that both transform the techniques within the sub-disciplines and enable the development of architectures that bring them together to add value to data. The commercial importance of the research area has been widely recognised. The VADA programme brings together university researchers with commercial partners who are in desperate need of a new generation of data management tools. They will be contributing to the programme by funding research staff and students, providing substantial amounts of staff time for research collaborations, supporting internships, hosting visitors, contributing challenging real-life case studies, sharing experiences, and participating in technical meetings. These partners are both developers of data management technologies (LogicBlox, Microsoft, Neo) and data user organisations in healthcare (The Christie), e-commerce (LambdaTek, PricePanda), finance (AllianceBernstein), social networks (Facebook), security (Horus), smart cities (FutureEverything), and telecommunications (Huawei).

View the original record at the funder ↗

Researchers

Alvaro Fernandes (Co-Investigator)Andreas Pieris (Co-Investigator)Dan Olteanu (Co-Investigator)Georg Gottlob (Principal Investigator)John Keane (Co-Investigator)Leonid Libkin (Co-Investigator)Norman Paton (Co-Investigator)Oscar Buneman (Co-Investigator)Paolo Guagliardo (Co-Investigator)Sebastian Maneth (Co-Investigator)Thomas Lukasiewicz (Co-Investigator)Wenfei Fan (Co-Investigator)

Related Research

Grants with similar aims, by meaning.

DIVA: Data Intensive Visual Analytics - Provenance and Uncertainty in Human Terrain Analysis
EPSRC Centre for Doctoral Training in Enhancing Human Interactions and Collaborations with Data and Intelligence Driven Systems
QuantiCode: Intelligent infrastructure for quantitative, coded longitudinal data
Cleaning Integrated Data: An Approach based on Conditional Constraints and Data Provenance
Smart Data Analytics for Business and Local Government

Original classification

Research Grant

Plain English summaries and category classifications on this site are generated by AI and may not perfectly reflect the original research.