MONAKA
About FUJITSU-MONAKA:
The HPC AI team at FRIPL is dedicated to advancing capabilities of AI and high-performance computing (HPC) for the next-generation ARM architecture-based 2nm CPU, FUJITSU-MONAKA. This processor, set to be released in 2027, is designed for use in data centers with a focus on AI-driven applications such as machine learning, deep learning, real-time big data processing, and large language models (LLMs).
Our team is at the forefront of optimizing AI frameworks to fully harness the power of FUJITSU-MONAKA, ensuring state-of-the-art performance and energy efficiency. We are deeply committed to sustainability, contributing to the realization of a carbon-neutral society through innovative, power-efficient computing solutions.
In collaboration with the open-source software community, we actively contribute to the development of AI-accelerated software stacks tailored for data-intensive workloads. This collaboration ensures that FUJITSU-MONAKA meets the demands of modern computing and aligns with Fujitsu’s vision of creating a more sustainable world through trust and innovation.
The HPC ML team at FRIPL is committed to pushing boundaries of computational innovation with ML algorithms. We focus on AI/ML framework engineering with advanced optimization for high-performance applications on FUJITSU-MONAKA. Our work spans a variety of cutting-edge technologies and methods to ensure that AI/ML framework engineering is efficient and scalable.
Our teams:
Our team’s expertise spans several key areas
Computational Data Science: Our research enhances computational performance of foundational frameworks, e.g., Scikit-Learn, XGBoost, Statsmodels (Time Series), OpenBLAS etc., to maximize efficiency for high-performance data science with large-scale data processing and predictive modelling.
Advanced Numerical Libraries: Optimizing complex computations with Scalable Vector Extensions (SVE), tailored for advanced ML applications ensuring high efficiency and reliability as required by modern AI models.
Scalability and Multithreading: Building scalable AI/ML software with efficient parallelism, reducing synchronization overhead, and enhancing workload distribution across multiple cores and multi-nodes. Our innovations with OpenMP, pthreads, TBB, etc. enable AI models to scale with high performance as data volumes and model complexities increase.
Key OSS Contributions & Optimizations
- PR#2614:
Enabled oneDAL and Scikit Learn Intelex on Arm using OpenBLAS as OSS reference backend for performance acceleration of ML workloads. This was one of the first contributions to UXL Foundation.
- PR#4741:
Contributed on Scalable AI with enhancing Core Utilization in OpenBLAS as Math Computational Library for AI/ML frameworks with Pthreads & OpenMP threading backend to improve workload distribution across multi-core systems.
- PR#2917:
Scalable Vector Extension (SVE) based tuning optimization enhances numerical computing performance for AI/ML models ensuring to fully leverage vectorized computations on Arm-based FUJITSU-MONAKA.
- PR#2807:
Development of SPBLAS (Sparse BLAS) and VSL (Vector Statistical Library) kernels to optimize sparse & vectorized numerical operations, accelerating core ML computations.
- PR#5091:
Small GEMM kernel tuning boosts the efficiency of matrix multiplications, improving execution times for HPC and AI/ML workloads.
A. Deep Learning Team
As artificial intelligence (AI) continues to drive innovation across various sectors, the role of deep learning has emerged as a cornerstone in this transformative era. HPC Deep Learning Team at FRIPL is at the forefront of technological innovation, tasked with the critical mission of advancing the capabilities and performance of deep learning frameworks for FUJITSU – MONAKA. Our ongoing research aims to accelerate the inference & fine-tuning performance of deep learning models on CPUs by efficient AI framework engineering.
Our team’s expertise spans several key areas:
- Foundational Frameworks:
Our research enhances computational performance of AI frameworks such as PyTorch, TensorFlow, JAX, ONNX etc. to maximize the performance for various Deep Learning applications at scale serving to solve complex problems around the globe.
- Advanced Compute Libraries:
Optimizing complex computations tailored for advanced DL Algorithms and models by leveraging advanced Arm supported Single Instruction Multiple Data (SIMD) such as Scalable Vector Extensions (SVE), ensuring high accuracy and reliability with enhanced performance.
- Scalability and Multithreading:
Building scalable AI Software with efficient parallelism, reducing synchronization overhead, and enhancing workload distribution across multiple cores and multi-node. Our focus is to work with OpenMP, TBB threading backends to enhance the scalability of DL workloads for high number of CPU Cores.
Key OSS Contributions & Optimizations:
- PR# 1818 :
Enabled BRGEMM based MatMul JIT kernels in oneDNN accelerating Deep Learning & Generative AI Workloads on Arm CPUs
- PR# 119571:
Extending the PyTorch vec backend for SVE (ARM). Initially only NEON vec backend was available for Arm CPUs, with this work latest SIMD has been implemented in PyTorch.
B. Data Platform Team
Data Platform team primarily focuses on enabling and optimizing various databases and distributed framework. Due to the fast development of open-source databases and its growing commercial popularity, our team focusses on various categories of data platforms like Relational databases, vector databases, NoSQL databases, graph databases, data lakes, large-scale real-time data processing & analytics engine, etc. to optimize for FUJITSU-MONAKA. Some of our major focus areas are described in the subsequent sections.
Our team’s expertise spans several key areas:
Loading component...
C. LLM Research Team
Loading component...
Our team’s expertise spans several key areas:
Loading component...
Key OSS Contributions & Optimizations:
