GitHub - shaunstanislauslau/pyBlazing: BlazingSQL is a lightweight, GPU accelerated, SQL engine built on RAPIDS.

A lightweight, GPU accelerated, SQL engine built on the RAPIDS.ai ecosystem.

BlazingSQL is a GPU accelerated SQL engine built on top of the RAPIDS ecosystem. RAPIDS is based on the Apache Arrow columnar memory format, and cuDF is a GPU DataFrame library for loading, joining, aggregating, filtering, and otherwise manipulating data.

BlazingSQL is a SQL interface for cuDF, with various features to support large scale data science workflows and enterprise datasets.

Query Data Stored Externally - a single line of code can register remote storage solutions, such as Amazon S3.
Simple SQL - incredibly easy to use, run a SQL query and the results are GPU DataFrames (GDFs).
Interoperable - GDFs are immediately accessible to any RAPIDS library for data science workloads.

Check out our 5-min quick start notebook using BlazingSQL.

Getting Started

Please reference our docs to find out how to install BlazingSQL.

Querying a CSV file in Amazon S3 with BlazingSQL:

For example:

from blazingsql import BlazingContext
bc = BlazingContext()

bc.s3('dir_name', bucket_name='bucket_name', access_key_id='access_key', secret_key='secret_key')

# Create Table from CSV
bc.create_table('taxi', '/dir_name/taxi.csv')

# Query
result = bc.sql('SELECT count(*) FROM taxi GROUP BY year(key)').get()
result_gdf = result.columns

#Print GDF 
print(result_gdf)

Examples

Getting Started Guide - Google Colab
Netflow Demo - Google Colab
Taxi cuML Linear Regression - Google Colab

Documentation

You can find our full documentation at the following site

Build/Install from Source

See build instructions.

Contributing

Have questions or feedback? Post a new github issue.

Please see our guide for contributing to BlazingSQL.

Contact

Feel free to join our Slack chat room: RAPIDS Slack Channel

You may also email us at info@blazingsql.com or find out more details on the BlazingSQL site

License

Apache License 2.0

RAPIDS AI - Open GPU Data Science

The RAPIDS suite of open source software libraries aim to enable execution of end-to-end data science and analytics pipelines entirely on GPUs. It relies on NVIDIA® CUDA® primitives for low-level compute optimization, but exposing that GPU parallelism and high-bandwidth memory speed through user-friendly Python interfaces.

Apache Arrow on GPU

The GPU version of Apache Arrow is a common API that enables efficient interchange of tabular data between processes running on the GPU. End-to-end computation on the GPU avoids unnecessary copying and converting of data off the GPU, reducing compute time and cost for high-performance analytics common in artificial intelligence workloads. As the name implies, cuDF uses the Apache Arrow columnar data format on the GPU. Currently, a subset of the features in Apache Arrow are supported.

Name		Name	Last commit message	Last commit date
Latest commit History 380 Commits
.github/ISSUE_TEMPLATE		.github/ISSUE_TEMPLATE
.settings		.settings
blazingsql		blazingsql
img		img
pyblazing		pyblazing
tests		tests
.gitignore		.gitignore
.project		.project
.pydevproject		.pydevproject
CONTRIBUTING.md		CONTRIBUTING.md
LICENSE		LICENSE
README.md		README.md
changelog		changelog
environment_pygdf.yml		environment_pygdf.yml
print_env.sh		print_env.sh
requirements.txt		requirements.txt
setup.py		setup.py

License

shaunstanislauslau/pyBlazing

Folders and files

Latest commit

History

Repository files navigation

Getting Started

Examples

Documentation

Build/Install from Source

Contributing

Contact

License

RAPIDS AI - Open GPU Data Science

Apache Arrow on GPU

About

Resources

License

Stars

Watchers

Forks

Languages