A highly practical approach toward achieving minimum data sets storage cost in the cloud

Date

2013

Authors

Yuan, D.
Yang, Y.
Liu, X.
Li, W.
Cui, L.
Xu, M.
Chen, J.

Editors

Advisors

Journal Title

Journal ISSN

Volume Title

Type:

Journal article

Citation

IEEE Transactions on Parallel and Distributed Systems, 2013; 24(6):1234-1244

Statement of Responsibility

Conference Name

Abstract

Massive computation power and storage capacity of cloud computing systems allow scientists to deploy computation and data intensive applications without infrastructure investment, where large application data sets can be stored in the cloud. Based on the pay-as-you-go model, storage strategies and benchmarking approaches have been developed for cost-effectively storing large volume of generated application data sets in the cloud. However, they are either insufficiently cost-effective for the storage or impractical to be used at runtime. In this paper, toward achieving the minimum cost benchmark, we propose a novel highly cost-effective and practical storage strategy that can automatically decide whether a generated data set should be stored or not at runtime in the cloud. The main focus of this strategy is the local-optimization for the tradeoff between computation and storage, while secondarily also taking users' (optional) preferences on storage into consideration. Both theoretical analysis and simulations conducted on general (random) data sets as well as specific real world applications with Amazon's cost model show that the cost-effectiveness of our strategy is close to or even the same as the minimum cost benchmark, and the efficiency is very high for practical runtime utilization in the cloud.

School/Discipline

Dissertation Note

Provenance

Description

Access Status

Rights

Copyright 2013 IEEE Transactions on Parallel and Distributed Systems

License

Grant ID

Call number

Persistent link to this record