Published June 2026 | Version v1
Dissertation Open

Improving Data's Utility

  • 1. University of Chicago

Contributors

Description

Users must make many different data management decisions about how to store, transform, and process their data in the cloud to best serve organizational goals. Because cloud services employ on-demand pricing models, dimensions such as monetary costs are just as important as task performance when evaluating the efficacy of a potential data management decision.

These problems each require different expertise and concern different data systems and are thus often solved in isolation using a mixture of heuristic signals, monitoring tools to detect issues, and manual investigation and tuning. We argue that across seemingly disparate data management problems, these approaches could be improved by identifying a common signal that directly indicates the benefit an organization achieves from its data. This benefit, called data utility, is a function of the value achieved from tasks run on data and the costs of running tasks and maintaining data. Though the particular utility function may vary between data managers and organizations, data's utility can be increased by reducing costs or improving the value achieved from tasks on that data.

This dissertation presents solutions that improve data's utility by building cheaper execution plans across multiple cloud data warehouses for analytical workloads and which provide users with an indicator for the value they have achieved from their data as a signal to make better data management decisions. These solutions together provide concrete mechanisms to enable users to directly interact with and improve their data's utility.

Files

Tapan_PhD_Dissertation.pdf

Files (1.7 MB)

Name Size Download all
md5:7fb6a0c47f6c624d96a8338ad0f5cb46
1.7 MB Preview Download

Additional details

Identifiers

Other
oai:uchicago.tind.io:17039

UChicago Information

Division(s)
Physical Sciences Division
Department(s)
Computer Science