diff --git a/02_activities/assignments/DC_Cohort/Assignment-1_Part1.pdf b/02_activities/assignments/DC_Cohort/Assignment-1_Part1.pdf new file mode 100644 index 000000000..29b1b5735 Binary files /dev/null and b/02_activities/assignments/DC_Cohort/Assignment-1_Part1.pdf differ diff --git a/02_activities/assignments/DC_Cohort/Assignment1.md b/02_activities/assignments/DC_Cohort/Assignment1.md index f650c9752..72f566aa7 100644 --- a/02_activities/assignments/DC_Cohort/Assignment1.md +++ b/02_activities/assignments/DC_Cohort/Assignment1.md @@ -210,4 +210,6 @@ Consider, for example, concepts of fariness, inequality, social structures, marg ``` Your thoughts... +Riz's story, as well as the reality faced by many people in Pakistan, seems so distant and shocking at first glance. However, upon some more reflection, I find there are aspects of this ethical discussion that I can relate to. As an international student since my undergraduate career, I've always been flagged when applying for scholarships, entering the country, and even exploring job opportunities. Although I understand the importance of differentiating residence/immigration status to ensure the safety and fairness in the country, at times, the fact that my identity and value all of a sudden gets undermined by one line in the country's database - "international student visa" - crushes my dreams and hopes I once brought with me when I first arrived in Canada. For example, because I immediately disquality for most scholarship or bursary opportunities as an international PhD student, I am unable to bring funds independently for my supervisor. Although unintended, this leads to a sense of inferiority within my lab and department, as the inability to bring in funds is frowned upon by supervisors. This immediate disqualitifaction because of my status in the databases despite my grades, experience, and outputs outperforming other candidates seem quite unfair. Although necessary from the perspective of security, marginalization and inequalities are discreetly implied by our governmental and even institutional databases. Another thought I had was the question of privacy with the emerging intersection of technology and our societies. When you travel to Asia (e.g., China), they now simply scan your face to enter the country or even to make purchases at convenience stores... Yes, this may seem convenient, but this also means it is that much easier for the government to follow and track your steps - perhaps there's a database of every action, purchase, or even word I've said! + ``` diff --git a/02_activities/assignments/DC_Cohort/Assignment2.md b/02_activities/assignments/DC_Cohort/Assignment2.md index 01f991d02..818f924fa 100644 --- a/02_activities/assignments/DC_Cohort/Assignment2.md +++ b/02_activities/assignments/DC_Cohort/Assignment2.md @@ -57,6 +57,11 @@ The store wants to keep customer addresses. Propose two architectures for the CU ``` Your answer... +Type 1 - Overwrite +The CUSTOMER_ADDRESS table has one row per customer, and when their address changes, the old address is simply overwritten. No history is kept. This type is simple to implement, but you permanently lose the old data. + +Type 2 - Retain changes +The table allows multiple rows per customer. Each row has additional columns like 'effective_date', 'expiry_date', and potentially an 'current' flag. So, when an address changes, the old row is essentially cl;osed (i.e., an expiry date is inputted, and the current flage = 0) and a new row is inserted. Full history can be preserved. ``` *** @@ -192,4 +197,9 @@ Consider, for example, concepts of labour, bias, LLM proliferation, moderating c ``` Your thoughts... +The article highlights how systems like ImageNet were built on the work of thousands of low-paid crowdworkers categorizing images, oftne under poor conditions for minimal pay. This raises serious concerns about fair compensation and credit. The outputs that emerge from this intense labour, sophisticated neural networks worth billions and fame for the spearheading professor, generate wealth tha almost never flows back to the poeple who made them possible. This also makes me reconsider the data cleaning and training processes that happen within our lab. To be frank, we often assign undergraduate students or even volunteers to spend hours a week to tag photos of food packages to identify elements of food marketing and labels, as well as flag outliers in our nutrient database, while we post-graduate students, post-docs, and professors, leverage these cleaned databases to run more 'complex' analyses that get published and recognized. However, in reality, it is because of the hard work and labour of our undergraduate students and volunteers that our analyses and work are even made possibe. We really need to do a better job recognizing and compensating their efforts. + +Another issue important to this story is the concept of embedded bias, which was also touched upon in last week's ethics writeup. Because humans are the ones labelling and tagging data, human prejudices can get embedded directly into models. If labellers consistently associate certain images, words, or concepts with particular groups, the model learns and amplifies those associations. And as thse models are deployed at massive scale, as we are seeing today, small biases become large societal level problems. For example, a hiring algorithm trained on biased data might reject thousands of qualified candidates; or a content moderation model trianed on subjective lables might silence certain communities. + +Perhaps the deepest issue is one of transparency and accountability ... or perhaps ignorance. Users almost blindly interact with AI systems these days as if they are neutral, objective tools. But they are anything but. They encode the priorities of whoever funded them, built them, and labelled their data. Without transparency from suppliers and intentionality of users to be educated about their choices, society will continue to run under these embedded biases that will eventually shape individuals' lives and society's values, cultures, and governance. ``` diff --git a/02_activities/assignments/DC_Cohort/assignment1.sql b/02_activities/assignments/DC_Cohort/assignment1.sql index 2ec561e2a..04f3d2181 100644 --- a/02_activities/assignments/DC_Cohort/assignment1.sql +++ b/02_activities/assignments/DC_Cohort/assignment1.sql @@ -6,7 +6,8 @@ --SELECT /* 1. Write a query that returns everything in the customer table. */ --QUERY 1 - +SELECT * +FROM customer; @@ -16,8 +17,10 @@ /* 2. Write a query that displays all of the columns and 10 rows from the customer table, sorted by customer_last_name, then customer_first_ name. */ --QUERY 2 - - +SELECT * +FROM customer +ORDER BY customer_last_name, customer_first_name +LIMIT 10; --END QUERY @@ -27,7 +30,10 @@ sorted by customer_last_name, then customer_first_ name. */ /* 1. Write a query that returns all customer purchases of product IDs 4 and 9. Limit to 25 rows of output. */ --QUERY 3 - +SELECT * +FROM customer_purchases +WHERE product_id IN (4,9) +LIMIT 25; @@ -42,7 +48,10 @@ filtered by customer IDs between 8 and 10 (inclusive) using either: Limit to 25 rows of output. */ --QUERY 4 - +SELECT *, (quantity * cost_to_customer_per_qty) AS price +FROM customer_purchases +WHERE customer_id BETWEEN 8 AND 10 +LIMIT 25; @@ -55,7 +64,12 @@ Using the product table, write a query that outputs the product_id and product_n columns and add a column called prod_qty_type_condensed that displays the word “unit” if the product_qty_type is “unit,” and otherwise displays the word “bulk.” */ --QUERY 5 +SELECT product_id, product_name +, CASE WHEN product_qty_type = 'unit' THEN 'unit' + ELSE 'bulk' + END as prod_qty_type_condensed +FROM product; @@ -66,7 +80,14 @@ if the product_qty_type is “unit,” and otherwise displays the word “bulk. add a column to the previous query called pepper_flag that outputs a 1 if the product_name contains the word “pepper” (regardless of capitalization), and otherwise outputs 0. */ --QUERY 6 - +SELECT product_id, product_name +, CASE WHEN product_qty_type = 'unit' THEN 'unit' + ELSE 'bulk' + END as prod_qty_type_condensed +, CASE WHEN product_name LIKE '%pepper%' THEN 1 + ELSE 0 + END as pepper_flag +FROM product; @@ -78,7 +99,13 @@ contains the word “pepper” (regardless of capitalization), and otherwise out vendor_id field they both have in common, and sorts the result by market_date, then vendor_name. Limit to 24 rows of output. */ --QUERY 7 +SELECT * +FROM vendor AS v +INNER JOIN vendor_booth_assignments as vba + ON v.vendor_id = vba.vendor_id +ORDER BY vba.market_date, v.vendor_name +LIMIT 24; @@ -92,8 +119,10 @@ Limit to 24 rows of output. */ /* 1. Write a query that determines how many times each vendor has rented a booth at the farmer’s market by counting the vendor booth assignments per vendor_id. */ --QUERY 8 - - +SELECT vendor_id, + COUNT(*) AS vendor_booth_assignments +FROM vendor_booth_assignments +GROUP BY vendor_id; --END QUERY @@ -105,8 +134,14 @@ of customers for them to give stickers to, sorted by last name, then first name. HINT: This query requires you to join two tables, use an aggregate function, and use the HAVING keyword. */ --QUERY 9 - - +SELECT c.customer_first_name +, c.customer_last_name +, SUM(cp.quantity *cp.cost_to_customer_per_qty) AS total_spent +FROM customer AS c +JOIN customer_purchases AS cp ON c.customer_id = cp.customer_id +GROUP BY c.customer_id +HAVING total_spent > 2000 +ORDER BY c.customer_last_name, c.customer_first_name; --END QUERY @@ -124,6 +159,11 @@ When inserting the new vendor, you need to appropriately align the columns to be VALUES(col1,col2,col3,col4,col5) */ --QUERY 10 +CREATE TEMP TABLE new_vendor AS +SELECT * FROM vendor; + +INSERT INTO new_vendor (vendor_id, vendor_name, vendor_type, vendor_owner_first_name, vendor_owner_last_name) +VALUES (10, 'Thomass Superfood Store', 'Fresh Focused', 'Thomas', 'Rosenthal'); @@ -138,7 +178,11 @@ HINT: you might need to search for strfrtime modifers sqlite on the web to know and year are! Limit to 25 rows of output. */ --QUERY 11 - +SELECT customer_id +, strftime ('%m', market_date) AS purchase_month +, strftime ('%Y', market_date) AS purchase_year +FROM customer_purchases +LIMIT 25; @@ -152,7 +196,12 @@ HINTS: you will need to AGGREGATE, GROUP BY, and filter... but remember, STRFTIME returns a STRING for your WHERE statement... AND be sure you remove the LIMIT from the previous query before aggregating!! */ --QUERY 12 - +SELECT customer_id +, SUM(quantity * cost_to_customer_per_qty) AS total_spent_april_2022 +FROM customer_purchases +WHERE strftime ('%m', market_date) = '04' + AND strftime ('%Y', market_date) = '2022' +GROUP BY customer_id; diff --git a/02_activities/assignments/DC_Cohort/assignment2-section1.drawio.png b/02_activities/assignments/DC_Cohort/assignment2-section1.drawio.png new file mode 100644 index 000000000..025d6e265 Binary files /dev/null and b/02_activities/assignments/DC_Cohort/assignment2-section1.drawio.png differ diff --git a/02_activities/assignments/DC_Cohort/assignment2.sql b/02_activities/assignments/DC_Cohort/assignment2.sql index f7515f625..763b73afb 100644 --- a/02_activities/assignments/DC_Cohort/assignment2.sql +++ b/02_activities/assignments/DC_Cohort/assignment2.sql @@ -23,7 +23,10 @@ Edit the appropriate columns -- you're making two edits -- and the NULL rows wil All the other rows will remain the same. */ --QUERY 1 +SELECT +product_name || ', ' || COALESCE(product_size, ' ')|| ' (' || COALESCE(product_qty_type, 'unit') || ')' +FROM product; --END QUERY @@ -41,8 +44,15 @@ HINT: One of these approaches uses ROW_NUMBER() and one uses DENSE_RANK(). Filter the visits to dates before April 29, 2022. */ --QUERY 2 +SELECT + customer_id + , market_date + , DENSE_RANK() OVER(PARTITION BY customer_id ORDER BY market_date) AS visit_number - +FROM customer_purchases +WHERE market_date < '2022-04-29' +GROUP BY customer_id, market_date +; --END QUERY @@ -53,7 +63,20 @@ only the customer’s most recent visit. HINT: Do not use the previous visit dates filter. */ --QUERY 3 +SELECT * + +FROM ( + SELECT + customer_id + , market_date + , DENSE_RANK() OVER (PARTITION BY customer_id ORDER BY market_date DESC) AS visit_number + + FROM customer_purchases + GROUP BY customer_id, market_date +) AS rev_ranked +WHERE visit_number = 1 +; --END QUERY @@ -66,8 +89,20 @@ You can make this a running count by including an ORDER BY within the PARTITION Filter the visits to dates before April 29, 2022. */ --QUERY 4 - - +SELECT + customer_id + , product_id + , market_date + , quantity + , cost_to_customer_per_qty + , COUNT(*) OVER( + PARTITION BY customer_id, product_id + ORDER BY market_date + ) AS times_purchased + +FROM customer_purchases +WHERE market_date < '2022-04-29' +; --END QUERY @@ -85,8 +120,16 @@ Remove any trailing or leading whitespaces. Don't just use a case statement for Hint: you might need to use INSTR(product_name,'-') to find the hyphens. INSTR will help split the column. */ --QUERY 5 +SELECT + product_name, + CASE + WHEN INSTR(product_name, '-') > 0 + THEN TRIM(SUBSTR(product_name, INSTR(product_name, '-') +1)) + ELSE NULL + END AS description - +FROM product +; --END QUERY @@ -94,8 +137,18 @@ Hint: you might need to use INSTR(product_name,'-') to find the hyphens. INSTR w /* 2. Filter the query to show any product_size value that contain a number with REGEXP. */ --QUERY 6 +SELECT + product_name, + product_size, + CASE + WHEN INSTR(product_name, '-') > 0 + THEN TRIM(SUBSTR(product_name, INSTR(product_name, '-') +1)) + ELSE NULL + END AS description - +FROM product +WHERE product_size REGEXP '[0-9]' +; --END QUERY @@ -111,8 +164,32 @@ HINT: There are a possibly a few ways to do this query, but if you're struggling with a UNION binding them. */ --QUERY 7 - - +WITH daily_sales AS( + SELECT + market_date, + ROUND(SUM(quantity * cost_to_customer_per_qty), 2) AS total_sales + FROM customer_purchases + GROUP BY market_date +), + +ranked_sales AS ( + SELECT + market_date, + total_sales, + RANK() OVER (ORDER BY total_sales DESC) AS best_rank, + RANK() OVER (ORDER BY total_sales ASC) AS worst_rank + FROM daily_sales +) + +SELECT market_date, total_sales, 'Best Day' AS label +FROM ranked_sales +WHERE best_rank = 1 + +UNION + +SELECT market_date, total_sales, 'Worst Day' AS label +FROM ranked_sales +WHERE worst_rank = 1; --END QUERY @@ -132,9 +209,24 @@ How many customers are there (y). Before your final group by you should have the product of those two queries (x*y). */ --QUERY 8 +SELECT + v.vendor_name, + p.product_name, + ROUND(5 * vi.original_price * COUNT(c.customer_id), 2) AS total_revenue +FROM( + SELECT DISTINCT vendor_id, product_id, original_price + FROM vendor_inventory -- 8 rows, so final output should be 8 rows too. +) AS vi +JOIN vendor v + ON vi.vendor_id = v.vendor_id +JOIN product p + ON vi.product_id = p.product_id +CROSS JOIN customer c +GROUP BY v.vendor_name, p.product_name +; --END QUERY @@ -144,8 +236,18 @@ This table will contain only products where the `product_qty_type = 'unit'`. It should use all of the columns from the product table, as well as a new column for the `CURRENT_TIMESTAMP`. Name the timestamp column `snapshot_timestamp`. */ --QUERY 9 - - +DROP TABLE IF EXISTS temp.product_units; +CREATE TEMP TABLE IF NOT EXISTS product_units AS +SELECT + product_id, + product_name, + product_size, + product_category_id, + product_qty_type, + CURRENT_TIMESTAMP AS snapshot_timestamp +FROM product +WHERE product_qty_type = 'unit' +; --END QUERY @@ -155,9 +257,11 @@ Name the timestamp column `snapshot_timestamp`. */ This can be any product you desire (e.g. add another record for Apple Pie). */ --QUERY 10 +INSERT INTO product_units +VALUES + (1, 'Apple Pie', 'whole', 1, 'unit', CURRENT_TIMESTAMP) - - +; --END QUERY @@ -167,9 +271,11 @@ This can be any product you desire (e.g. add another record for Apple Pie). */ HINT: If you don't specify a WHERE clause, you are going to have a bad time.*/ --QUERY 11 +DELETE FROM product_units +WHERE product_id = 7 +; - - +SELECT * FROM product_units --END QUERY @@ -191,8 +297,21 @@ Finally, make sure you have a WHERE statement to update the right row, When you have all of these components, you can run the update statement. */ --QUERY 12 +ALTER TABLE product_units +ADD current_quantity INT; - +UPDATE product_units +SET current_quantity = COALESCE( + ( + SELECT vi.quantity + FROM vendor_inventory vi + WHERE vi.product_id = product_units.product_id + ORDER BY vi.market_date DESC + LIMIT 1 + ), 0 +) +; +SELECT * FROM product_units --END QUERY