Professional Documents
Culture Documents
738 PDFsam PythonNotesForProfessionals
738 PDFsam PythonNotesForProfessionals
We assume that a customer can have n orders, an order can have m items, and items can be ordered more
multiple times
orders_df = pd.DataFrame()
orders_df['customer_id'] = [1,1,1,1,1,2,2,3,3,3,3,3]
orders_df['order_id'] = [1,1,1,2,2,3,3,4,5,6,6,6]
orders_df['item'] = ['apples', 'chocolate', 'chocolate', 'coffee', 'coffee', 'apples',
'bananas', 'coffee', 'milkshake', 'chocolate', 'strawberry', 'strawberry']
.
.
Now, we will use pandas transform function to count the number of orders per customer
# First, we define the function that will be applied per customer_id
count_number_of_orders = lambda x: len(x.unique())
# And now, we can transform each group using the logic defined above
orders_df['number_of_orders_per_cient'] = ( # Put the results into a new column that
is called 'number_of_orders_per_cient'
orders_df # Take the original dataframe
.groupby(['customer_id'])['order_id'] # Create a separate group for each
customer_id & select the order_id
.transform(count_number_of_orders)) # Apply the function to each group
separately