The Big Data era is upon us: data is being generated, collected and analyzed at an unprecedented scale, and data-driven decision making is sweeping through all aspects of technology and society. Since the value of data increases exponentially when it can be linked and fused with other data, addressing the big data integration challenge is critical to realizing the promise of Big Data -- and conversely, Big Data techniques are critical to the goals of simplifying data integration.
The convergence of Big Data and data integration is emerging in many forms, largely motivated by the goals of integrating structured data on the Web or across communities. Increasingly we are seeing problems where (i) the number of data sources, even for a single domain, has grown to be in the tens of thousands, (ii) many of the data sources are very dynamic, as large volumes of newly collected data are continuously made available, (iii) the data sources are extremely heterogeneous in their structure, with considerable variety even for conceptually similar entities, and (iv) the data sources are of widely differing quality, with significant differences in the coverage, accuracy and timeliness of data provided.
This workshop will focus on the progress that is being made to address
these novel challenges faced by big data integration. It will bring
together researchers from data integration, data cleaning, machine
learning, and data analysis to address these issues, and we expect to
identify a range of open problems for the community. Questions of
interest during the workshop include:
When confronted with large numbers of available data sources and a higher variety of data, how can we resolve heterogeneity of the data at the schema level?
With the higher velocity, greater variety, and significant differences in the veracity of data, how can we adapt existing entity resolution techniques to resolve heterogeneity of the data at the instance level? How do we carry such information forward in our query processing model?
How can we improve scalability of data fusion, which aims at resolving conflicts from different sources and finding the truths that reflect the real world?
How can we apply emerging techniques for learning concepts, properties, and ontologies from the Web into the data integration framework, and provide utility?
What are other open problems or promising new techniques for big data integration, such as integrating crowdsourced data, integrating data from data markets, providing an exploration tool for data sources, and so on?