vault backup: 2026-06-04 14:22:26
This commit is contained in:
1 parent
e7e1b545ae
commit
69d7a58a2c
1 file changed
+4
-9
+4
-9
@@ -1,17 +1,12 @@
|
|||||||
#rs/class/ad450 #rs/discussion
|
#rs/class/ad450 #rs/discussion
|
||||||
- - -
|
- - -
|
||||||
1. **Personal Experience with Data Visualization:**
|
1. **Personal Experience with Data Visualization:**
|
||||||
- Describe an instance where you have used data visualization in a project or a task. What was the purpose of the visualization, and what tools did you use to create it?
|
- One project that I have used data visualization for is a machine learning project for one of my other classes. I talked about this project in a previous discussion and there were several places where I used data visualization to get a better idea of the data I had and use it to make better decisions on preprocessing. One of the visualizations that I used was a pairplot which shows scatter plots for each combination of features in the dataset. It will also color the dots by the output class which can be very helpful to see which features might have correlations that could help to predict the output class or that might interfere with predictions. When used in combination with a correlation heat map between all the features correlation can be found. Using this I was able to identify several features that were essentially duplicates and would not provide any more information to the model. To create these visualizations I used matplotlib and seaborn. These libraries are very helpful for creating visualizations of data, especially when also using pandas.
|
||||||
- One project that I have used data visualization for is a machine learning project for one of my other classes. I talked about this project in a previous discussion and there were several places where I used data visualization to get a better idea of the data I had and use it to make better decisions on preprocessing. One of the visualizations that I used was a pairplot which shows scatter plots for each combination of features in the dataset. It will also color the dots by the output class which can be very helpful to see which features might have correlations that could help to predict the output class or that might interfere with predictions. When used in combination with a correlation heatmap between all the features correlation can be found. Using this I was able to identify several features that were essentially duplicates and would not provide any more information to the model. To create these visualizations I used matplotlib and seaborn. These libraries are very helpful for creating visualizations of data, especially when also using pandas.
|
|
||||||
2. **Challenges and Learning:**
|
2. **Challenges and Learning:**
|
||||||
- What challenges did you encounter while working on your data visualization? How did you overcome these challenges? Share any learning points or insights you gained from this experience.
|
|
||||||
- One of the challenges that I encountered while creating the pairplot visualization was the amount of data that I was displaying. The dataset that I was using has over 100k instances and since the pairplot was graphing a scatterplot between each pair of features it would take a very long time to generate. To make it faster to create I randomly sampled a much smaller number of the instances to still get a good idea of the distribution and relationships between the features while reducing computation time.
|
- One of the challenges that I encountered while creating the pairplot visualization was the amount of data that I was displaying. The dataset that I was using has over 100k instances and since the pairplot was graphing a scatterplot between each pair of features it would take a very long time to generate. To make it faster to create I randomly sampled a much smaller number of the instances to still get a good idea of the distribution and relationships between the features while reducing computation time.
|
||||||
3. **Impact of Visualization on Understanding:**
|
3. **Impact of Visualization on Understanding:**
|
||||||
- Reflect on how the visualization helped in better understanding or communicating the data. Was there a notable difference in comprehension or decision-making due to the visualization?
|
- Visualizing the correlation between different features in my dataset helped to inform many of the preprocessing choices that I made. There were several features that I dropped from using in the model due to very strong correlations with other features or a lack of usefulness. Another visualization that I used was a bar plot of the distribution of the output class which showed how imbalanced the data was. This helped me to make the decision to use SMOTE to balance the data for use in the model to help improve the accuracy and effectiveness of predictions.
|
||||||
- Visualizing the correlation between different features in my dataset helped to inform many of the preprocessing choices that I made. There were several features that I dropped from using in the model due to very strong correlations with other features or a lack of usefulness. Another visualization that I used was a barplot of the distribution of the output class which showed how imbalanced the data was. This helped me to make the decision to use SMOTE to balance the data for use in the model to help improve the accuracy and effectiveness of predictions.
|
|
||||||
4. **Tool Exploration:**
|
4. **Tool Exploration:**
|
||||||
- Have you experimented with any data visualization tools (like Tableau, PowerBI, Microsoft Excel, Google Sheets, RStudio, Jupyter Notebook etc.)? Describe your experience with these tools. Which one did you find most effective, and why?
|
- The data visualization tools that I have used include Excel, Google Sheets and Jupyter Notebook (matplotlib/seaborn). Excel and Google Sheets can be good if you have to make a quick visualization of data without having to write any code and are easier to understand and use. They aren't as customizable as something like matplotlib though and have a lot of limitations in the types of visualizations that they can create. They are good at making scatter, line and bar plots but for anything else it is probably better to use a more dedicated library.
|
||||||
- The data visualization tools that I have used include Excel, Google Sheets and Jupyter Notebook (matplotlib/seaborn). Excel and Google Sheets can be good if you have to make a quick visualization of data without having to write any code and are easier to understand and use. They aren't as customizable as something like matplotlib though and have a lot of limitations in the types of visualizations that they can create. They are good at making scatter, line and barplots but for anything else it is probably better to use a more dedicated library.
|
|
||||||
5. **Future Applications:**
|
5. **Future Applications:**
|
||||||
- How do you envision applying data visualization techniques in your future projects or career? Discuss any specific areas where you think data visualization can be particularly impactful.
|
- Data visualization can be particularly impactful when you need to communicate results to people who have not worked on the project that you are doing. Showing results and graphs visually is a great way to help people better understand the data, especially if they don't have a technical background. Data visualization can also be very helpful to see visual relationships between data that might not be as obvious from statistics and metrics alone. In the future I will likely keep using data visualization techniques to see relationships between features in my data and to ensure that I understand the data that I am using fully. Plotting things like cluster assignments or showing the data in 2D with PCA can be very helpful to gain more insight into how models are performing and why they might be making the decisions they are.
|
||||||
- Data visualization can be particularly impactful when you need to communicate results to people who have not worked on the project that you are doing.
|
|
||||||
Reference in new issue
Block a user