Repository navigation
07 Pandas Integration
You can see simpler examples of using pandas.DataFrame.read_table with Chemaxon molecules in the Calculators and Molecular similiarity notebook. But here are some additional examples showing how to work with Chemaxon molecules in pandas.DataFrames.
By default, molecules are being printed in the cxsmiles format. In case of issues with that, it is written in the cxsmarts format. If none of them are possible for some reason, representation falls back to use the original format, which was recognized during the import process.
import pandas as pd
from chemaxon.io import import_mol, open_for_import
mol = import_mol('CC(=O)NC1=CC=C(O)C=C1')
mol_lst = []
with open_for_import('../mol_w_sgroups.mol') as mol_importer:
for m in mol_importer:
mol_lst.append(m)
d = {'molecule': [mol] + mol_lst }
df_mols = pd.DataFrame(data=d)
df_mols| molecule | |
|---|---|
| 0 | CC(=O)NC1=CC=C(O)C=C1 |c:9,t:4,6| |
| 1 | CC(C)CC1=CC=CC=C1 |c:6,8,t:4| |
We also provide a utility function that returns a dictionary object that already contains the Molecule objects, so you can easily create a DataFrame from it.
from chemaxon.pandasutil import load_molecules_for_pandas, prepare_molecules_for_pandas
df_mols_2 = pd.DataFrame(data=load_molecules_for_pandas('../mol_w_sgroups.mol'))
df_mols_2| Molecule | |
|---|---|
| 0 | CC(C)CC1=CC=CC=C1 |c:6,8,t:4| |
Utility method prepare_molecules_for_pandas enables creating a DataFrame object from a list of chemaxon.Molecule objects. read_properties_to_columns parameter facilitates the creation of columns from sdf molecule properties.
from chemaxon.pandasutil import prepare_molecules_for_pandas
mol_lst_with_properties = []
with open_for_import('../mol_with_properties.sdf') as mol_importer:
for m in mol_importer:
mol_lst_with_properties.append(m)
df_mols_with_properties = pd.DataFrame(data=prepare_molecules_for_pandas(mol_lst_with_properties, read_properties_to_columns=True))
df_mols_with_properties| Molecule | test_boolean_property | test_int_array_property | test_str_property | test_int_property | test_double_property | test_double_array_property | |
|---|---|---|---|---|---|---|---|
| 0 | C1=CC=C(C=C1)C1=CC=C(C=C1)C1=CC=CC=C1 |c:0,2,4... | True | [1, 2, 3, 4, 5] | asd | 42 | 4.2 | [1.2, 3.4, 5.22, 6.7889] |
In case of creating HTML output from a pandas.DataFrame, you can use the helper function mol_to_svg_formatter to visualize the molecules as SVG images.
from chemaxon.pandasutil import mol_to_svg_formatter
df_mols.to_html('../web_view.html', escape=False, formatters=dict(molecule=mol_to_svg_formatter))HTML output combined with molecules loaded from an input file. File load is parametrized in order to see the exported molecule in the cxsmiles column:
from chemaxon.pandasutil import mol_to_svg_formatter
df_mol_with_custom_columns = pd.DataFrame(data=load_molecules_for_pandas(file_path='../mol_w_sgroups.mol', mol_obj_column="mols", mol_str_column="cxsmiles"))
df_mol_with_custom_columns.to_html('../web_view_cxsmiles_column.html', escape=False, formatters=dict(mols=mol_to_svg_formatter))Since the Molecule objects are being stored in the DataFrame, not just their representation, you can easily calculate properties for them and store the results in new columns.
from chemaxon.calculations import logp
df_mols['LogP'] = df_mols['molecule'].apply(lambda m: logp(m))
df_mols| molecule | LogP | |
|---|---|---|
| 0 | CC(=O)NC1=CC=C(O)C=C1 |c:9,t:4,6| | 0.92 |
| 1 | CC(C)CC1=CC=CC=C1 |c:6,8,t:4| | 3.71 |
You can also easily create new molecule columns based on existing columns, that contain molecules in any of the supported formats.
d = {'SMILES': ['CN1C=NC2=C1C(=O)N(C)C(=O)N2C'], 'name': ['coffein'] }
df = pd.DataFrame(data=d)
df = pd.concat([df, pd.DataFrame.from_records([{'SMILES' : 'CC[C@H](C)[C@@H]1NC(=O)[C@H](CC2=CC=C(O)C=C2)NC(=O)[C@@H](N)CSSC[C@H](NC(=O)[C@H](CC(N)=O)NC(=O)[C@H](CCC(N)=O)NC1=O)C(=O)N1CCC[C@H]1C(=O)N[C@@H](CC(C)C)C(=O)NCC(N)=O', 'name' : 'oxytocin'}])])
df['molecule'] = df['SMILES'].apply(lambda s: import_mol(s))
df| SMILES | name | molecule | |
|---|---|---|---|
| 0 | CN1C=NC2=C1C(=O)N(C)C(=O)N2C | coffein | CN1C=NC2=C1C(=O)N(C)C(=O)N2C |c:2,4| |
| 0 | CC[C@H](C)[C@@H]1NC(=O)[C@H](CC2=CC=C(O)C=C2)N... | oxytocin | CC[C@H](C)[C@@H]1NC(=O)[C@H](CC2=CC=C(O)C=C2)N... |
You can also use molecule properties in DataFrame objects:
from chemaxon.io import open_for_import
with open_for_import('../mol_with_properties.sdf') as mol_iterator:
mols = list(mol_iterator)
d = {'molecule': mols }
df_props = pd.DataFrame(data=d)
df_props['string property'] = df_props['molecule'].apply(lambda m: m.get_property('test_str_property'))
df_props['int array property'] = df_props['molecule'].apply(lambda m: m.get_property('test_int_array_property'))
df_props| molecule | string property | int array property | |
|---|---|---|---|
| 0 | C1=CC=C(C=C1)C1=CC=C(C=C1)C1=CC=CC=C1 |c:0,2,4... | asd | [1, 2, 3, 4, 5] |