Cylon

Cylon is a fast, scalable distributed memory data parallel library for processing structured data. Cylon implements a set of relational operators to process data. While ”Core Cylon” is implemented using system level C/C++, multiple language interfaces (Python and Java ) are provided to seamlessly integrate with existing applications, enabling both data and AI/ML engineers to invoke data processing operators in a familiar programming language. By default it works with MPI for distributing the applications.

Internally Cylon uses Apache Arrow to represent the data in a column format.

The documentation can be found at https://cylondata.org

Email - cylondata@googlegroups.com

Mailing List - Join

Getting Started

We can use Conda to install PyCylon. At the moment Cylon only works on Linux Systems. The Conda binaries need Ubuntu 16.04 or higher.

conda create -n cylon-0.4.0 -c cylondata pycylon python=3.7
conda activate cylon-0.4.0

Now lets run our first Cylon application inside the Conda environment. The following code creates two DataFrames and joins them.

from pycylon import DataFrame, CylonEnv
from pycylon.net import MPIConfig

df1 = DataFrame([[1, 2, 3], [2, 3, 4]])
df2 = DataFrame([[1, 1, 1], [2, 3, 4]])

# local merge
df3 = df1.merge(right=df2, on=[0, 1])
print("Local Merge")
print(df3)

Now lets run a parallel version of this program. Here if we create n processes (parallelism), n instances of the program will run. They will each load two DataFrames in their memory and do a distributed join among the DataFrames. The results will be created in the parallel processes as well.

from pycylon import DataFrame, CylonEnv
from pycylon.net import MPIConfig
import random

# distributed join
env = CylonEnv(config=MPIConfig())

df1 = DataFrame([random.sample(range(10*env.rank, 15*(env.rank+1)), 5),
                 random.sample(range(10*env.rank, 15*(env.rank+1)), 5)])
df2 = DataFrame([random.sample(range(10*env.rank, 15*(env.rank+1)), 5),
                 random.sample(range(10*env.rank, 15*(env.rank+1)), 5)])
df2.set_index([0], inplace=True)
print("Distributed Join")
df3 = df1.join(other=df2, on=[0], env=env)
print(df3)

You can run the above program in the Conda environment by using the following command. It uses mpirun command with 2 parallel processes.

mpirun -np 2 python <name of your python file>

Compiling Cylon

Refer to the documentation on how to compile Cylon

Compiling on Linux

License

Cylon uses the Apache Lincense Version 2.0

Name	Name	Last commit message	Last commit date
Latest commit arupcsedu [Radical-Cylon] Summit build and test (#685 ) Dec 5, 2023 03e0338 · Dec 5, 2023 History 1,307 Commits
.github/workflows	.github/workflows	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
aws	aws	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
conda	conda	Ucc ucx redis local ucx (#673 )	Sep 16, 2023
cpp	cpp	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
data	data	Fixing 554 (#558 )	Jan 5, 2022
docker	docker	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
docs	docs	Cylon Release Version 0.6.0 (#650 )	Mar 7, 2023
java	java	Ucc integration (#591 )	Jul 14, 2022
python	python	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
rivanna	rivanna	[Radical-Cylon] Summit build and test (#685 )	Dec 5, 2023
summit	summit	[Radical-Cylon] Summit build and test (#685 )	Dec 5, 2023
target	target	[Radical-Cylon] Summit build and test (#685 )	Dec 5, 2023
.gitignore	.gitignore	Adding to docker docs (#498 )	Sep 21, 2021
.gitmodules	.gitmodules	adding cylonflow as a submodule (#593 )	Jul 8, 2022
.travis.yml	.travis.yml	UCX integration (#439 )	Jun 20, 2021
LICENSE	LICENSE	adding the folder structure for java	Feb 10, 2020
NOTICE.txt	NOTICE.txt	adding third party flat_hash_map (#268 )	Dec 29, 2020
README.md	README.md	Update README.md	Apr 26, 2021
build.py	build.py	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
build.sh	build.sh	Ucc ucx redis rivanna updates (#681 )	Nov 25, 2023
requirements.txt	requirements.txt	[Cylon-RP] Cylon scaling test with radical-pilot (#661 )	Aug 3, 2023

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Cylon

Getting Started

Compiling Cylon

License

About

Releases 8

Packages

Contributors 19

Languages

License

cylondata/cylon

Folders and files

Latest commit

History

Repository files navigation

Cylon

Getting Started

Compiling Cylon

License

About

Topics

Resources

License

Stars

Watchers

Forks

Releases 8

Packages 0

Contributors 19

Languages

Packages