vault backup: 2026-06-04 14:12:21

This commit is contained in:
ben committed 2026-06-04 14:12:21 -07:00
1 parent 5a0ceea830
commit e7e1b545ae
1 file changed
+4 -2
@@ -8,8 +8,10 @@
- One of the challenges that I encountered while creating the pairplot visualization was the amount of data that I was displaying. The dataset that I was using has over 100k instances and since the pairplot was graphing a scatterplot between each pair of features it would take a very long time to generate. To make it faster to create I randomly sampled a much smaller number of the instances to still get a good idea of the distribution and relationships between the features while reducing computation time. - One of the challenges that I encountered while creating the pairplot visualization was the amount of data that I was displaying. The dataset that I was using has over 100k instances and since the pairplot was graphing a scatterplot between each pair of features it would take a very long time to generate. To make it faster to create I randomly sampled a much smaller number of the instances to still get a good idea of the distribution and relationships between the features while reducing computation time.
3. **Impact of Visualization on Understanding:** 3. **Impact of Visualization on Understanding:**
- Reflect on how the visualization helped in better understanding or communicating the data. Was there a notable difference in comprehension or decision-making due to the visualization? - Reflect on how the visualization helped in better understanding or communicating the data. Was there a notable difference in comprehension or decision-making due to the visualization?
- - Visualizing the correlation between different features in my dataset helped to inform many of the preprocessing choices that I made. There were several features that I dropped from using in the model due to very strong correlations with other features or a lack of usefulness. Another visualization that I used was a barplot of the distribution of the output class which showed how imbalanced the data was. This helped me to make the decision to use SMOTE to balance the data for use in the model to help improve the accuracy and effectiveness of predictions.
4. **Tool Exploration:** 4. **Tool Exploration:**
- Have you experimented with any data visualization tools (like Tableau, PowerBI, Microsoft Excel, Google Sheets, RStudio, Jupyter Notebook etc.)? Describe your experience with these tools. Which one did you find most effective, and why? - Have you experimented with any data visualization tools (like Tableau, PowerBI, Microsoft Excel, Google Sheets, RStudio, Jupyter Notebook etc.)? Describe your experience with these tools. Which one did you find most effective, and why?
- The data visualization tools that I have used include Excel, Google Sheets and Jupyter Notebook (matplotlib/seaborn). Excel and Google Sheets can be good if you have to make a quick visualization of data without having to write any code and are easier to understand and use. They aren't as customizable as something like matplotlib though and have a lot of limitations in the types of visualizations that they can create. They are good at making scatter, line and barplots but for anything else it is probably better to use a more dedicated library.
5. **Future Applications:** 5. **Future Applications:**
- How do you envision applying data visualization techniques in your future projects or career? Discuss any specific areas where you think data visualization can be particularly impactful. - How do you envision applying data visualization techniques in your future projects or career? Discuss any specific areas where you think data visualization can be particularly impactful.
- Data visualization can be particularly impactful when you need to communicate results to people who have not worked on the project that you are doing.