Microsoft Releases Open Source Distributed Machine Learning Library
Written by Kay Ewbank   
Tuesday, 04 January 2022

Microsoft has released an open-source library for creating massively scalable machine learning (ML) pipelines. SynapseML was until now known as MMLSpark, and it unifies several existing ML frameworks and new Microsoft algorithms in a single, scalable API that’s usable across Python, R, Scala, and Java.

SynapseML builds on Apache Spark and SparkML, adding deep learning and data science tools to the Spark ecosystem.


It integrates Spark Machine Learning pipelines with the Open Neural Network Exchange (ONNX), LightGBM, The Cognitive Services, Vowpal Wabbit, and OpenCV, so through these tools provides highly-scalable predictive and analytical models for a variety of data sources. SynapseML also includes the HTTP on Spark project, meaning users can embed web services into their SparkML models.


SynapseML simplifies this experience by unifying many different ML learning frameworks with a single API that is scalable, data- and language-agnostic, and that works for batch, streaming, and serving applications. It’s designed to help developers focus on the high-level structure of their data and tasks, not the implementation details and idiosyncrasies of different ML ecosystems and databases.

The unified API provides a standard way to use the tools, so developers can make use of multiple ML frameworks where necessary. It can also train and evaluate models on single-node, multi-node, and elastically resizable clusters of computers, making it simpler to scale up as required.

Describing the new software, Mark Hamilton, a Microsoft software engineer, said that many tools in SynapseML don’t require a large labelled training dataset. Instead, SynapseML provides simple APIs for pre-built intelligent services, such as Azure Cognitive Services, to quickly solve large-scale AI challenges related to both business and research.

SynapseML lets developers embed 45 different ML services directly into their systems and databases. The latest release includes added support for distributed form recognition, conversation transcription, and translation. These ready-to-use algorithms can parse a wide variety of documents, transcribe multi-speaker conversations in real time, and translate text to over 100 different languages.

SynapseML also extends Spark's Structured Streaming engine, meaning that jobs that run on the Structured Streaming engine can be used via a web service.

SynapseML is available now.


More Information

SynapseML On GitHub

SynapseML Website

Related Articles

Microsoft Open Sources Natural Language Processing Tool

More AI Tools From Microsoft

Apache Ignite Adds Spark DataFrames Support

.NET For Apache Spark Updated

To be informed about new articles on I Programmer, sign up for our weekly newsletter, subscribe to the RSS feed and follow us on Twitter, Facebook or Linkedin.


GitHub Sees Exponential Rise In AI

Developers are flocking to AI creating an explosion of generative AI activity in open source. The 11th annual Octoverse report, unveiled at last week's GitHub Universe event recorded 65K public g [ ... ]

Grafana Adds New Tools

Grafana Labs has announced new tools to make it easier to analyze application data on Grafana Cloud. The announcements are an Application Observability tool for Grafana Cloud, and Grafana Beyla, the e [ ... ]

More News




or email your comment to: