initial commit
This commit is contained in:
commit
fbd12d4a8e
320 files changed
+124298
No files matched your search
+95
@@ -0,0 +1,95 @@
|
||||
#rs/assignment #rs/class/ad450
|
||||
- - -
|
||||
## Task 1
|
||||
```SQL
|
||||
SELECT country, "1990" + "1991" + "1992" + "1993" + "1994" AS total_volume FROM coffee_export
|
||||
ORDER BY total_volume DESC
|
||||
```
|
||||
![[Week4Task1.csv]]
|
||||
Looking at the data we can see that countries like Brazil and Columbia are some of the largest exporters of coffee between 1990 and 1994. This likely means that they have large industries and are great countries to look to for production since coffee production is more common in them.
|
||||
## Task 2
|
||||
```SQL
|
||||
SELECT RANK() OVER (ORDER BY b.total_volume DESC) AS rank, b.country, b.total_volume FROM (
|
||||
SELECT
|
||||
country,
|
||||
"1999_2000" + "2000_2001" + "2001_2002" + "2002_2003" + "2003_2004"
|
||||
AS total_volume FROM coffee_production
|
||||
) b
|
||||
LIMIT 10
|
||||
```
|
||||
![[Week4Task2.csv]]
|
||||
Alternate for task 2
|
||||
```SQL
|
||||
WITH summed_years as (
|
||||
SELECT
|
||||
country,
|
||||
"1999_2000" + "2000_2001" + "2001_2002" + "2002_2003" + "2003_2004" AS total_volume
|
||||
FROM coffee_production
|
||||
)
|
||||
|
||||
SELECT RANK() OVER (ORDER BY total_volume DESC) AS rank, country, total_volume FROM summed_years
|
||||
LIMIT 10
|
||||
```
|
||||
|
||||
There are a lot of the same countries on the top exporters list on the top producers list meaning that their coffee industries are likely largely focused on international markers rather than domestic ones. Sourcing all product from one country could be a risk since if something happened to their coffee industry or economy it would be challenging to pivot somewhere else for production.
|
||||
## Task 3
|
||||
```SQL
|
||||
WITH summed_years as (
|
||||
SELECT
|
||||
country,
|
||||
coffee_type,
|
||||
"1990_1991" + "1991_1992" + "1992_1993" + "1993_1994" AS total_volume
|
||||
FROM coffee_production
|
||||
), ranked as (
|
||||
SELECT
|
||||
DENSE_RANK() OVER (PARTITION BY coffee_type ORDER BY total_volume DESC) AS rank,
|
||||
country,
|
||||
coffee_type,
|
||||
total_volume
|
||||
FROM summed_years
|
||||
)
|
||||
|
||||
SELECT coffee_type, country, total_volume from ranked
|
||||
WHERE rank = 2
|
||||
```
|
||||
"Runner-up" countries such as these could present an opportunity since they likely still have established coffee industries but there might not be as much competition from other large companies. It could be easier to expand and grow without as much competition from other large brands.
|
||||
## Task 4
|
||||
```SQL
|
||||
WITH summed_exports as (
|
||||
SELECT
|
||||
country,
|
||||
"1995" + "1996" + "1997" + "1998" + "1999" + "2000" AS total_export
|
||||
FROM coffee_re_export --looking at countries that are re-exporting coffee rather than exporting for the first time
|
||||
), summed_imports AS (
|
||||
SELECT
|
||||
country,
|
||||
"1995" + "1996" + "1997" + "1998" + "1999" + "2000" AS total_import
|
||||
FROM coffee_import
|
||||
)
|
||||
|
||||
SELECT
|
||||
COALESCE(e.country, i.country) AS country,
|
||||
COALESCE(e.total_export, 0) AS export_vol,
|
||||
COALESCE(i.total_import, 0) AS import_vol,
|
||||
COALESCE(e.total_export, 0) + COALESCE(i.total_import, 0) as total_vol
|
||||
FROM summed_exports e
|
||||
FULL JOIN summed_imports i ON e.country = i.country
|
||||
ORDER BY total_vol DESC
|
||||
LIMIT 5
|
||||
```
|
||||
Countries such as the United States, Germany and France act as some of the world's primary "coffee clearinghouses" since they import and then re-export the largest quantities of coffee. This means that they are focusing more on processing more than production.
|
||||
## Task 5
|
||||
```SQL
|
||||
SELECT
|
||||
i.country AS importing_country,
|
||||
i.total_import AS importing_amount,
|
||||
e.country AS exporting_country,
|
||||
e.total_export AS exporting_amount
|
||||
FROM coffee_import i
|
||||
CROSS JOIN LATERAL (
|
||||
SELECT e.country, e.total_export FROM coffee_export e
|
||||
ORDER BY ABS(i.total_import - e.total_export) ASC
|
||||
LIMIT 1
|
||||
) e
|
||||
```
|
||||
With the data we have it is not possible to show the country of origin for the coffee. The closest I got was guessing based on matching similar values of imports and exports but in most cases countries will be importing from or exporting to multiple sources. Additional data in the form of percentage import or export from each country or origin of export and import would be needed to complete this request.
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
#rs/discussion #rs/class/ad450
|
||||
- - -
|
||||
"[T]heir greatest opportunity to add value is not in creating reports or presentations for
|
||||
senior executives but in innovating with customer-facing products and processes."
|
||||
|
||||
This quote stood out to me because it often seems like a lot of work with data is behind the scenes of companies and that it is mostly used for internal analytics. While this is all used to create better products for the user it doesn't seem like data science is often used to directly innovate on customer-facing products. In the cases where this does happen it seems like there can be a large impact to improve a product for the consumer with more data and analysis.
|
||||
|
||||
"There simply aren’t a lot of people with their combination of scientific
|
||||
background and computational and analytical skills."
|
||||
|
||||
This quote stood out to me since it was a little surprising that there are not enough people for this job. The article was written in 2012 though, so this field was still relatively new and there was not the amount of people in computer science and related fields that there are now. I wonder if there is still a shortage or with how many people are in computer science there is enough supply for the demand of data scientist roles.
|
||||
+20
@@ -0,0 +1,20 @@
|
||||
#rs/discussion #rs/class/ad450
|
||||
- - -
|
||||
- **Ask Phase Reflection:**
|
||||
- Describe a time when you had to define a problem or understand stakeholder expectations in a project or study. How did you approach this task, and what challenges did you face?
|
||||
- One time that I had to define a problem and use the data analysis process was when I worked on creating an analysis tool for my high school robotics team. There is a website that publishes data about competitions and team performance and I wanted to use this data to create custom metrics about team performance and do some analysis. This project was largely for myself but I set the goal of creating a metric that could perform similarly to some of the others that exist such as offensive power rating or estimated points added.
|
||||
- **Prepare Phase Experience:**
|
||||
- Share an instance where you had to collect and prepare data for analysis. What types of data did you use, and how did you ensure its relevance and objectivity?
|
||||
- The data that I used came from thebluealliance.com which is a site that publishes all match data from First Robotics competitions. This data is published directly from competitions and is the primary source of match data for teams and districts. While it can be inaccurate sometimes the inaccuracies come from how the matches were scored at events and not the data entered into the system and the website will always reflect match results.
|
||||
- **Process Phase Challenges:**
|
||||
- Reflect on a situation where you had to clean or process data. What difficulties did you encounter, and how did you overcome them?
|
||||
- I pulled the data from an api so it was pretty well structured. The processing that I had to do on the data was largely ensuring that I was pulling from the correct matches and events rather than having to clean the data.
|
||||
- **Analyze Phase Insights:**
|
||||
- Discuss your experience with analyzing data. What tools or methods did you use, and how did you interpret the results?
|
||||
- One of the things that I did to analyze the data was to try and calculate a strength of schedule metric for teams at an event based on who they faced in their qualification matches. I did this by looking at how strong the teams that they were facing were and used data from after the event like match scores to give a metric of how challenging their schedule was. This wasn't the most helpful since it could only be used after an event but in the future I might expand it to predict the strength of a schedule before the matches have happened.
|
||||
- **Share Phase Application:**
|
||||
- Recall a time when you had to present your data findings. How did you communicate your insights, and what techniques did you use to make your presentation effective?
|
||||
- I didn't present much of my findings since I didn't get to a point where I was finished but if I did I would likely give examples and show some manual verification to communicate accuracy. I would also probably go through my methods for calculating the metric to show others what I was doing and how I was getting the numbers I was getting.
|
||||
- **Act Phase Impact:**
|
||||
- Think about an occasion where your data analysis led to actionable insights. What were the outcomes, and how did your analysis influence decision-making?
|
||||
- I didn't apply my findings much but the data could lead to better strategy decisions at or after competitions since you can look at the strength of the schedule of other teams at the event and determine if they might be ranked higher than they should be or lower than they should be. This is important in the review phase for an event or when making decisions of what teams to pick at the event.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
#rs/discussion #rs/class/ad450
|
||||
- - -
|
||||
- **Planning Stage:**
|
||||
- If I was planning a survey for a school project I would consider what I was looking for and what data I wanted to gather. Once I define the goal of the project I can start thinking about what questions I might want to ask or what pieces of data to gather. Having a clear idea of the goal of the survey is important to ensure that all the needed data is gathered.
|
||||
- **Capturing Data:**
|
||||
- One of the sources of data that I would likely use the most is surveys since you can get qualitative and quantitative data from them and they are pretty easy to use to gather information. There can be many biases with this approach though so other sources based on the project would be useful as well. This could include the usage or activity of a certain service.
|
||||
- **Managing Data:**
|
||||
- I have a place to organize notes and writing for all of my classes and other projects and I have faced a few challenges in organizing it. One of these is the many different topics of my writing and currently I have everything organized into folders and subfolders based on the topic and subtopics that it relates to. This makes it easier to find information although other approaches like including tags could be useful in the future as well.
|
||||
- **Analyzing Data:**
|
||||
- One time where I had to analyze information to make a decision was when I was building my new computer. When doing this I went through many different sources including articles, videos, guides, and benchmarking sites in order to find the best parts for me. I created lists and compared the features, draw backs and pricing of many different components in order to find the best options and make a data-driven decision.
|
||||
- **Archiving Data:**
|
||||
- Every once in a while I will go through my phone and take photos off to free up space. I will often decide what to keep on my phone based on what I think I will want locally and how much space it is taking up. I will then move the photos and videos that I don't want to keep on my phone to a larger storage drive that I can still access although not quite as easily. This allows me to keep older photos while still keeping the number locally on my phone low.
|
||||
- **Destroying Data:**
|
||||
- I'm sure that there are many protocols for the deletion of sensitive information to ensure that it is not able to be recovered and I would make sure to follow the policy and procedure of where the information came from or where I might work.
|
||||
+14
@@ -0,0 +1,14 @@
|
||||
#rs/class/ad450 #rs/discussion
|
||||
- - -
|
||||
### Description
|
||||
For another one of my classes I recently completed a project where I had to analyze a dataset including cleaning the data before using it with machine learning models to predict one of the features. The dataset that I chose to work with was a collection of data from individual laps of Formula 1 races over the last 4 years. This was a very large dataset with over 100k instances and there were several things that I had do do to preprocess and clean the data.
|
||||
### Challenges
|
||||
- **Missing Values:** One of the features in the dataset had some instances with missing values. There were 66 of the rows of the feature describing the tire compound used by the team that were missing values. Without doing something to remove the missing data later steps in the process would fail since they don't have ways to deal with null values.
|
||||
- **Outliers:** A few of the features in the dataset (especially the LapTime feature) had some extreme outliers. Most of the values of the LapTime feature were between 0 and 200 seconds while there were a few that were over 2000 seconds. This could cause models later on to be biased and metrics such as average could also not properly represent the data.
|
||||
- **Redundant Features:** There were some features that were redundant and contained effectively the same data in the dataset. One of these was RaceProgress and LapNumber which are highly correlated and contain the same information in slightly different forms. This introduces duplicate data that could harm model performance later.
|
||||
### Strategies Used
|
||||
- **Missing Values:** Since there were only 66 missing values for the Compound feature I decided to drop the rows with the missing values. With over 100k rows 66 is a very small percentage of the data and dropping the rows is easier than imputing the data. This was effective since I didn't lose much of the data and didn't introduce the possible inaccuracy of imputation.
|
||||
- **Outliers:** To remove the outliers from the dataset I created a function that looked for values outside of a +/- 3 standard deviation range. This range contains 99.7% of the data so it will only remove values that are very extreme and biasing the data. This removed the outliers and made the data much more consistent. This was mostly effective but one of the features that I used the function on removed almost 1k instances. I decided to keep this but that probably means that the data was more spread rather than having outliers.
|
||||
- **Redundant Features:** For the redundant features I decided to remove one of them to get rid of the duplicate data. For RaceProgress and LapNumber I removed the LapNumber feature since RaceProgress was more granular and was there for likely more accurate. This was pretty effective although I could have performed some more analysis to better decide which feature was better to keep.
|
||||
### Reflections
|
||||
Overall I learned a lot about data cleaning with this project and this was one of the first times that I went through the full data pipeline. I especially learned about different considerations you need to keep in mind when preprocessing data for machine learning pipelines such as looking for outliers and removing or imputing missing data. In the future I don't think I would approach data cleaning very differently but I do think I have a better understanding and will likely be able to go more in depth in the future.
|
||||
+32
@@ -0,0 +1,32 @@
|
||||
#rs/discussion #rs/class/ad450
|
||||
- - -
|
||||
- **Role and Impact in Business**
|
||||
- Discuss the contribution of data engineers to the success of a business. How do they enable data-driven decision-making?
|
||||
- Data engineers contribute to the success of a business by giving access to important data about what is happening. Without data about what users want, their activity, and the success of the product it is very challenging for a business to grow, improve and create better products for its customers. Providing data allows for making decisions that align with what is happening and not just based on what seems correct or what some people think might be best. It can also help to measure success quantitatively to ensure that something is succeeding or to figure out when it is not.
|
||||
- **Essential Skills and Tools**
|
||||
- Identify and explain the most crucial skills for a data engineer. Why are these skills important, especially in the context of big data technologies and ETL processes?
|
||||
- Some of the most crucial skills for data engineers are being able to use languages such as Python, R and SQL as well as understanding different pipelines, databases and storage schemes. Being able to work with different programming languages is important because it unlocks many powerful tools for quickly and efficiently collecting, cleaning, and organizing data. Understanding different pipelines and schemes is also important so that data scientists are able to use what is best for the company they work at.
|
||||
- **Data Engineers vs. Data Scientists**
|
||||
- Compare and contrast the roles of data engineers and data scientists within a data analytics team. How do these roles complement each other?
|
||||
- Data engineers focus more on gathering data and ensuring it is ready for analysis and able to be gathered and analyzed. Data scientists focus on analyzing the data and looking for patterns significance and future trends. These two roles compliment each other by providing the two large steps of the process that are collecting and processing the data and analyzing it for use.
|
||||
- **Industry-Specific Challenges**
|
||||
- Examine the unique challenges data engineers might face in industries like healthcare, retail, and financial services. How do these challenges impact their work?
|
||||
- Industries such as these are more data intensive and require the processing and analysis for large quantities of data. In order to effectively operate companies in these industries must use data to impact their decision making so that they can make the best decisions to be the most profitable that they can be. The large quantities of data needed can make the job of data scientist more challenging since larger datasets can be more unwieldy and require more time for cleaning and preparing the data.
|
||||
- **Evolution of the Field**
|
||||
- Analyze how the field of data engineering has evolved and predict future trends. Consider the impact of emerging technologies and data volume growth.
|
||||
- The field of data engineering has become much more important for companies since using data to analyze trends and consumer behavior is vital for companies to stay competitive. Because of this data engineering has become a much more sought after position and as the role has also expanded to meet demands it has also split into many different more specialized roles. Along with more course and education around these topics the role has expanded and specialized to provide companies with the data they need. In the future this role will likely continue to expand, especially with companies continuing to collect more data on consumers and try to predict their behavior more.
|
||||
- **Career Pathways**
|
||||
- Discuss the educational and experiential pathways beneficial for aspiring data engineers. Assess the value of certifications, degrees, and hands-on learning.
|
||||
- There are several things beneficial for aspiring data engineers such as university degrees, projects and certifications. University degrees can include those in applied mathematics and computer science. Projects can also useful to build a portfolio and show potential employers what you can do and what you have worked on. Certificates can also be useful to show employers specific skill sets that you have learned and are proficient in. All of these options can be helpful to improve your skills and be more competitive in the job market.
|
||||
- **Ethical Considerations**
|
||||
- Explore the ethical considerations in data engineering, focusing on data privacy and security. How should data engineers approach these issues responsibly?
|
||||
- Privacy and security are important considerations for data engineers especially as more and more data is harvested from users by large companies. There likely isn't much data engineers can do at most companies since they are likely not the ones making the decisions about what data to use and record but ensuring that they follow the terms of service or privacy policy is important. Securely storing private information is also vital so that it is not leaked unintentionally.
|
||||
- **Real-World Applications**
|
||||
- Provide examples or case studies where data engineering played a critical role. Discuss the solutions implemented and their effectiveness.
|
||||
- One example that was discussed in one of the articles was about LinkedIn. Near the start the platform would just let users find others and create connections on their own. Someone had the idea for the platform to suggest others to users to increase the number of connections that they would make and used data on the platform to suggest users. These suggestions massively increased engagements and connection on the platform since it exposed users to others that they might not have connected with using data engineering and analysis.
|
||||
- **Collaboration with IT Professionals**
|
||||
- Describe how data engineers collaborate with other IT professionals. Highlight the importance of teamwork and communication in successful data projects.
|
||||
- Data engineering is not the only aspect of most projects and people to help with networking, UI design, backend development and more are necessary for projects to succeed. Communication is important to ensure that all of these different aspects work together and form one cohesive product or system.
|
||||
- **Future in AI and Machine Learning**
|
||||
- Discuss the evolving role of data engineers in the context of AI and machine learning. How must data engineers adapt to these technological changes?
|
||||
- Data engineers are necessary in the development of AI and machine learning systems since large amounts of data are required for them to be developed and function. These systems can also provide new tools for analyzing and organizing data and help to make the work of data engineers more efficient. Learning to effectively use these new tools is important to stay competitive and successful as a data engineer.
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
#rs/discussion #rs/class/ad450
|
||||
- - -
|
||||
"A decade later, the job is more in demand than ever with employers and recruiters."
|
||||
|
||||
After the last article I was wondering if the supply for data scientists had now risen to meet or surpass the demand and it seems like this is not the case. I expected that the amount of new people going into computer science careers would have increased the number of data scientists. Maybe the role of data scientist is still something that is not very known about so most people just try to become software engineers instead.
|
||||
|
||||
"Now, however, there has been a proliferation of related jobs to handle many of those tasks, including machine learning engineer, data engineer, AI specialist, analytics and AI translators, and data oriented product managers."
|
||||
|
||||
This quote stood out to me since it was interesting how much the role has split up in the last ten years. It makes sense that as the role of data scientist has expanded and become more important it has split up into many different roles to increase productivity but I didn't realize how many. This could also be a reason for the demand since there are many more job titles and positions that are required at large companies.
|
||||
Reference in new issue
Block a user