DESIGN AND OPTIMIZATION OF CLOUD-NATIVE DATA PROCESSING PIPELINES USING AUTOML FOR DYNAMIC WORKLOAD ADAPTATION AND COST MINIMIZATION

Authors

  • Nawal El Saadwi Nivil, Independent Researcher Author

Keywords:

Cloud-Native, AutoML, Data Pipelines, Cost Minimization, Dynamic Workload, Kubernetes, Resource Optimization, Serverless, Workflow Orchestration

Abstract

The rising demand for scalable and efficient data processing has driven the adoption of cloud-native architectures. However, designing pipelines that adapt automatically to fluctuating workloads while minimizing cost remains a complex challenge. This study proposes an AutoML-powered framework for dynamically optimizing data processing pipelines in cloud-native environments. The system automates configuration tuning, workload scaling, and cost optimization across compute resources. Our results demonstrate significant improvements in cost-efficiency and processing latency across varied workload patterns when compared with static rule-based systems

 

 

References

Liu, Y., et al. (2022). Resource-aware Scheduling in Cloud Pipelines via AutoML. ACM Transactions on Cloud Computing.

Sankaranarayanan, S. (2025). The Role of Data Engineering in Enabling Real-Time Analytics and Decision-Making Across Heterogeneous Data Sources in Cloud-Native Environments. International Journal of Advanced Research in Cyber Security (IJARC), 6(1), January-June 2025.

Adapa, C.S.R. (2025). Building a standout portfolio in master data management (MDM) and data engineering. International Research Journal of Modernization in Engineering Technology and Science, 7(3), 8082–8099. https://doi.org/10.56726/IRJMETS70424

Mukesh, V. (2024). A Comprehensive Review of Advanced Machine Learning Techniques for Enhancing Cybersecurity in Blockchain Networks. ISCSITR-International Journal of Artificial Intelligence, 5(1), 1–6.

Zhen, L., et al. (2021). Cost-Effective Optimization of Cloud Functions Using Bayesian AutoML. IEEE Transactions on Services Computing.

Chen, T., et al. (2020). A Survey on AutoML. ACM Computing Surveys, 54(8), 1–36.

Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource Management with Deep Reinforcement Learning. HotNets.

Adapa, C.S.R. (2025). Transforming quality management with AI/ML and MDM integration: A LabCorp case study. International Journal on Science and Technology (IJSAT), 16(1), 1–12.

S.Sankara Narayanan and M.Ramakrishnan, Software As A Service: MRI Cloud Automated Brain MRI Segmentation And Quantification Web Services, International Journal of Computer Engineering & Technology, 8(2), 2017, pp. 38–48.

Kumar, A., et al. (2020). Serverless Data Pipelines: Challenges and Opportunities. Proceedings of VLDB.

Sankar Narayanan .S, System Analyst, Anna University Coimbatore , 2010. INTELLECTUAL PROPERY RIGHTS: ECONOMY Vs SCIENCE &TECHNOLOGY. International Journal of Intellectual Property Rights (IJIPR) .Volume:1,Issue:1,Pages:6-10.

Google Cloud. (2021). Kubernetes AutoPilot Documentation.

Microsoft Azure. (2022). Azure Functions Premium Plan Scaling Guide.

Wang, S., et al. (2021). Dynamic Batch Scheduling for Serverless Workflows. IEEE CLOUD.

Yu, X., & Venkatraman, S. (2021). On-demand Scaling for Streaming Pipelines. IEEE ICDE.

Mukesh, V. (2025). Architecting intelligent systems with integration technologies to enable seamless automation in distributed cloud environments. International Journal of Advanced Research in Cloud Computing (IJARCC), 6(1),5-10.

Chandra Sekhara Reddy Adapa. (2025). Blockchain-Based Master Data Management: A Revolutionary Approach to Data Security and Integrity. International Journal of Information Technology and Management Information Systems (IJITMIS), 16(2), 1061-1076.

Li, C., et al. (2020). Kubeflow Pipelines for End-to-End ML Automation. Google Research.

Zhang, J., et al. (2022). Intelligent Orchestration in Cloud Workflows. IEEE Transactions on Network and Service Management.

Mukesh, V. (2022). Evaluating Blockchain Based Identity Management Systems for Secure Digital Transformation. International Journal of Computer Science and Engineering (ISCSITR-IJCSE), 3(1), 1–5.

Ghosh, R., & Shenoy, P. (2019). AutoScale: Dynamic Scaling in Kubernetes Clusters. ACM SoCC.

Sankar Narayanan .S System Analyst, Anna University Coimbatore , 2010. PATTERN BASED SOFTWARE PATENT.International Journal of Computer Engineering and Technology (IJCET) -Volume:1,Issue:1,Pages:8-17.

Yu, Q., et al. (2020). SLA-aware Cloud Task Placement using Bayesian Optimization. Journal of Cloud Computing.

Adapa, C.S.R. (2025). Cloud-based master data management: Transforming enterprise data strategy. International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 11(2), 1057–1065. https://doi.org/10.32628/CSEIT25112436

Lin, S., et al. (2021). Learning to Tune Cloud Infrastructure Configurations. NeurIPS Workshop on ML for Systems.

Reddi, V. J., et al. (2020). MLPerf Inference Benchmarking. IEEE Micro, 40(2), 8–16.

Xie, C., et al. (2021). FunctionFusion: Cost-Efficient Serverless Function Aggregation. USENIX ATC.

Mukesh, V., Joel, D., Balaji, V. M., Tamilpriyan, R., & Yogesh Pandian, S. (2024). Data management and creation of routes for automated vehicles in smart city. International Journal of Computer Engineering and Technology (IJCET), 15(36), 2119–2150. doi: https://doi.org/10.5281/zenodo.14993009

Zhao, Z., et al. (2020). Elasticity Optimization in Container-based Stream Processing. ACM DEBS.

Downloads

Published

2025-05-05

How to Cite

Nawal El Saadwi Nivil,. (2025). DESIGN AND OPTIMIZATION OF CLOUD-NATIVE DATA PROCESSING PIPELINES USING AUTOML FOR DYNAMIC WORKLOAD ADAPTATION AND COST MINIMIZATION. International Journal of Computer Science and Engineering Research and Development (IJCSERD), 15(3), 14-21. https://ijcserd.in/index.php/home/article/view/IJCSERD_15_03_003