Research
Fifteen years on the performance management of data centre workloads, from application servers to the edge.
Background
My research was about the performance management of data centre workloads: deciding where work runs, and when, so that shared infrastructure meets its goals. It began with my PhD on heterogeneous workload management in clouds, moved to the autonomic placement of mixed batch and transactional workloads with IBM Research, and to deadline-aware scheduling for MapReduce.
At the Barcelona Supercomputing Center I led the Data-Centric Computing research group, and with it we widened that work to big data, IoT streams, GPUs and non-volatile memory, and to AI for resource management: learning models, including conditional restricted Boltzmann machines, that predict how workloads behave and steer data centre operations. With industrial partners including IBM, Microsoft, Intel and Cisco, the same question reached fog and edge computing, where infrastructure is distributed and heterogeneous. Nearby Computing grew out of that last step.
Themes
AI for resource management
Machine learning models of workloads, from classical methods to deep learning and conditional restricted Boltzmann machines, to predict demand and guide placement and data centre operations.
Workload placement and scheduling
Holistic optimisation of software-defined data centres, with performance models that span heterogeneous infrastructure and workloads.
Fog and edge computing
Bridging cloud and edge for NFV and 5G, and the architectures that let cities and networks run applications close to where data is produced.
Big data cost-effectiveness
How configuration choices drive the runtime and price of Hadoop and Spark deployments. The group built the ALOJA open benchmarking platform.
IoT stream processing
Real-time composition, transformation and filtering of data streams. The group built the servIoTicy platform.
Storage and acceleration
Non-volatile memory, GPUs and FPGAs for data-centric and IO-bound applications.
Selected publications
AI and GPU management, service placement
- 2022Burst-aware predictive autoscaling for containerized microservices
- 2017Topology-aware GPU scheduling for learning workloads in cloud environments
- 2024Dexter: a performance-cost efficient resource allocation manager for serverless data analytics
AI for data centre operations
- 2020Adaptive sliding windows for improved estimation of data center resource utilization
- 2019Adaptive prediction models for data center resources utilization estimation
- 2020A highly parameterizable framework for Conditional Restricted Boltzmann Machine based workloads accelerated with FPGAs and OpenCL
- 2018Automatic generation of workload profiles using unsupervised learning pipelines
Edge and fog computing
- 2022Autonomous lifecycle management for resource-efficient workload orchestration for green edge computing
- 2017A new era for cities with fog computing
- 2017The unavoidable convergence of NFV, 5G, and fog: a model-driven approach to bridge cloud and edge
- 2022Automatic distributed deep learning using resource-constrained edge devices
Big data
- 2010Performance-driven task co-scheduling for MapReduce environments
- 2011Resource-aware adaptive scheduling for MapReduce clusters
- 2014ALOJA: a systematic study of Hadoop deployment variables to enable automated characterization of cost-effectiveness
More than 100 peer-reviewed papers in all. The full list is on Google Scholar and DBLP.