This notebook presents an empirical investigation into the 2023 Amazon Books dataset. By isolating the "Mystery" category, we model the probability of market success through granular metadata analysis, transitioning from raw data to actionable statistical insights.
The research utilizes metadata from the McAuley Lab (UCSD), processing millions of entries to identify key performance indicators. The data provides a high-fidelity look at pricing, product formats, and internal Kindle metadata features.
Establishes a curated sample space (S) defined by products with 500+ ratings and a 4.3+ star average.
Calculates the probability of Kindle, Paperback, and Hardcover editions as mutually exclusive market events.
Uses set theory and Venn visualizations to map the intersection of non-mutually exclusive category labels.
Analyzes digital-only features like X-Ray and Enhanced Typesetting to identify success correlations.
The analytical framework is designed to handle multi-label classification and overlapping variables. This allows for a deeper understanding of how "Mystery" sub-genres interact with physical and digital format preferences.