Menu
R to python data wrangling snippets
The
dplyr
package in R makes data wrangling significantly easier.The beauty of dplyr
is that, by design, the options available are limited.Specifically, a set of key verbs form the core of the package.Using these verbs you can solve a wide range of data problems effectively in a shorter timeframe.Whilse transitioning to Python I have greatly missed the ease with which I can think through and solve problems using dplyr in R.The purpose of this document is to demonstrate how to execute the key dplyr verbs when manipulating data using Python (with the pandas
package).dplyr is organised around six key verbs:
R to Python: Data wrangling with dplyr and pandas Raw. R-to-python-data-wrangling-basics.md R to python data wrangling snippets. The dplyr package in R makes data wrangling significantly easier. The beauty of dplyr is that, by design, the options available are limited. Specifically, a set of key verbs form the core of the package.
filter
: subset a dataframe according to condition(s) in a variable(s)select
: choose a specific variable or set of variablesarrange
: order dataframe by index or variablegroup_by
: create a grouped dataframesummarise
: reduce variable to summary variable (e.g. mean)mutate
: transform dataframe by adding new variables
The excellent pandas package in Python easily allows you to implement all of these actions (and much, much more!). Below are some snippets to highlight some of the more basic conversions.
June 8th 2018 update:All of this code still works in pandas and should ease the transition from R, but for those interested in getting the most out of the package I strongly recommend this series on modern pandas https://tomaugspurger.github.io/modern-1-intro.html
Filter
R
Python
Select
R
Python
Arrange
R
Python
Grouping
R
Python
Summarise / Aggregate df by group
R
Python
Mutate / transform df by group
R
Python
Distinct
R
Python
Sample
R
Python