GitHub - Le-Zheng/BigDL: BigDL: Distributed Deep Learning Library for Apache Spark

Distributed Deep Learning Library for Apache Spark

AI for Big Data

The AI for Big Data community includes the following projects:

BigDL: distributed deep learning library for Apache Spark
Analytics Zoo: distributed Tensorflow, PyTorch and Ray on Apache Spark (as well as Spark ML pipeline for BigDL)

What is BigDL?

BigDL is a distributed deep learning library for Apache Spark; with BigDL, users can write their deep learning applications as standard Spark programs, which can directly run on top of existing Spark or Hadoop clusters.

Rich deep learning support. Modeled after Torch, BigDL provides comprehensive support for deep learning, including numeric computing (via Tensor) and high level neural networks; in addition, users can load pre-trained Caffe or Torch models into Spark programs using BigDL.
Extremely high performance. To achieve high performance, BigDL uses Intel oneMKL, oneDNN and multi-threaded programming in each Spark task. Consequently, it is orders of magnitude faster than out-of-box open source Caffe or Torch on a single-node Xeon (i.e., comparable with mainstream GPU).
Efficiently scale-out. BigDL can efficiently scale out to perform data analytics at "Big Data scale", by leveraging Apache Spark (a lightning fast distributed data processing framework), as well as efficient implementations of synchronous SGD and all-reduce communications on Spark.

Why BigDL?

You may want to write your deep learning programs using BigDL if:

You want to analyze a large amount of data on the same Big Data (Hadoop/Spark) cluster where the data are stored (in, say, HDFS, HBase, Hive, Parquet, etc.).
You want to add deep learning functionalities (either training or prediction) to your Big Data (Spark) programs and/or workflow.
You want to leverage existing Hadoop/Spark clusters to run your deep learning applications, which can be then dynamically shared with other workloads (e.g., ETL, data warehouse, feature engineering, classical machine learning, graph analytics, etc.)

How to use BigDL?

It is highly recommended to use the high-level APIs provided by Analytics Zoo, including:

Spark ML pipeline support for BigDL
Keras-like API for BigDL

For additional information, you may refer to:

Citing BigDL

If you've found BigDL useful for your project, you may cite the paper as follows:

@inproceedings{SOCC2019_BIGDL,
  title={BigDL: A Distributed Deep Learning Framework for Big Data},
  author={Dai, Jason (Jinquan) and Wang, Yiheng and Qiu, Xin and Ding, Ding and Zhang, Yao and Wang, Yanzhang and Jia, Xianyan and Zhang, Li (Cherry) and Wan, Yan and Li, Zhichao and Wang, Jiao and Huang, Shengsheng and Wu, Zhongyuan and Wang, Yang and Yang, Yuhao and She, Bowen and Shi, Dongjie and Lu, Qi and Huang, Kai and Song, Guoqiong},
  booktitle={Proceedings of the ACM Symposium on Cloud Computing},
  publisher={Association for Computing Machinery},
  pages={50--60},
  year={2019},
  series={SoCC'19},
  doi={10.1145/3357223.3362707},
  url={https://arxiv.org/pdf/1804.05839.pdf}
}

Name		Name	Last commit message	Last commit date
Latest commit History 2,683 Commits
.github		.github
core @ 95eeaaf		core @ 95eeaaf
docker		docker
docs		docs
pyspark		pyspark
scripts		scripts
spark		spark
.gitignore		.gitignore
.gitmodules		.gitmodules
.travis.yml		.travis.yml
LICENSE		LICENSE
README.md		README.md
make-dist.sh		make-dist.sh
pom.xml		pom.xml
scalastyle_config.xml		scalastyle_config.xml

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

AI for Big Data

What is BigDL?

Why BigDL?

How to use BigDL?

Citing BigDL

About

Releases

Packages

Languages

License

Le-Zheng/BigDL

Folders and files

Latest commit

History

Repository files navigation

AI for Big Data

What is BigDL?

Why BigDL?

How to use BigDL?

Citing BigDL

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages