SIM-PIPE DryRunner: An approach for testing container-based big data pipelines and generating simulation data

Big data pipelines are becoming increasingly vital in a wide range of data intensive application domains such as digital healthcare, telecommunication, and manufacturing for efficiently processing data. Data pipelines in such domains are complex and dynamic and involve a number of data processing steps that are deployed on heterogeneous computing resources under the realm of the Edge-Cloud paradigm. The processes of testing and simulating big data pipelines on heterogeneous resources need to be able to accurately represent this complexity. However, since big data processing is heavily resource-intensive, it makes testing and simulation based on historical execution data impractical. In this paper, we introduce the SIM - PIPE Dry Runner approach - a dry run approach that deploys a big data pipeline step by step in an isolated environment and executes it with sample data; this approach could be used for testing big data pipelines and realising practical simulations using existing simulators.

Utgiver

Institute of Electrical and Electronics Engineers (IEEE)

Serie

IEEE Annual International Computer Software and Applications Conference (COMPSAC);2022 IEEE 46th Annual Computers, Software, and Applications Conference (COMPSAC)

Tidsskrift

IEEE Annual International Computer Software and Applications Conference (COMPSAC)