{"id":386,"date":"2025-04-09T16:01:38","date_gmt":"2025-04-09T16:01:38","guid":{"rendered":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/11691-zhiyang-wang\/"},"modified":"2025-06-19T23:00:26","modified_gmt":"2025-06-19T23:00:26","slug":"11691-zhiyang-wang","status":"publish","type":"post","link":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/11691-zhiyang-wang\/","title":{"rendered":"LLM Assisted Visualization Analysis Pipeline"},"content":{"rendered":"<h2>Project Description<\/h2>\n<p>\n    This project presents a scalable, end-to-end system for extracting, classifying, and interactively exploring large collections of scholarly figures using a multimodal language model (LLM). Leveraging Playwright automation and OpenAlex harvesting, we collected over 11,000 publication PDFs. Figures are isolated using pdffigures2 and a Faster R-CNN-based detector, followed by zero-shot chart-type classification via GPT-4o prompting. The resulting metadata populates a dual-mode exploration interface\u2014a traditional 2D dashboard and a 3D free-exploration environment built with Observable Framework, Three.js, and D3.js.<\/p>\n<p>Designed to support visualization service providers, the system enables rapid trend discovery, technique identification, and cross-domain consulting. Initial evaluations show a manual-verification accuracy of 91.2%, with the potential to reduce manual annotation efforts by over 98%. Future work will integrate Vision Transformer embeddings with t-SNE\/UMAP dimensionality reduction for taxonomy-free exploration.<\/p>\n<p>This work would not have been possible without the support of the ITP community, as well as Carolina Roe-Raymond (Princeton University) and Devin Richard Bayly (University of Arizona).  <\/p>\n<h2>Technical Details<\/h2>\n<p>\n    Data Acquisition &#038; Preprocessing:<\/p>\n<p>    1. OpenAlex API harvesting<\/p>\n<p>    2. Playwright browser automation<\/p>\n<p>    3. pdffigures2 for figure\u2013caption extraction<\/p>\n<p>    4. VisImages-Detection (Faster R\u2011CNN) for subfigure isolation<\/p>\n<p>Classification:<\/p>\n<p>    1. Zero-shot chart-type inference via OpenAI GPT\u20114o (version: 2024\u201108\u201106)<\/p>\n<p>    2. Caption-based prompt engineering<\/p>\n<p>Interface:<\/p>\n<p>    1. 2D dashboards: Observable Framework<\/p>\n<p>    2. 3D exploration: Three.js and D3.js<\/p>\n<p>    3. Future work: Taxonomy-free embedding using t\u2011SNE\/UMAP on Vision Transformer features  <\/p>\n<h2>Research\/Context<\/h2>\n<p>\n    This work proposes a quantitative analysis pipeline for visualization datasets, combining bibliometric and document analysis foundations. Using this pipeline, we built one of the largest institutional collections of visualization figures (>30k images), supporting flexible, user-defined taxonomy. Unlike prior projects such as Beagle\u2019s web extraction [Battle et al. 2018] and Vis30k\u2019s curated datasets [Chen et al. 2021], our approach eliminates manual annotation overhead by incorporating LLM-based zero-shot labeling. Drawing interface inspirations from Google\u2019s t-SNE Map [Diagne et al. 2018] and Duhaime\u2019s Three.js guide [Duhaime 2017], we designed a 3D interactive exploration environment. Comparisons to VisImages [Deng et al. 2022] emphasize challenges in multi-label figure handling and motivate subfigure isolation strategies. This work demonstrates how AI4Vis techniques can streamline visualization consulting, trend analysis, and technique discovery at scale across institutions.  <\/p>\n","protected":false},"excerpt":{"rendered":"<p>A scalable multimodal LLM\u2011powered pipeline that automates the extraction, classification, and interactive 2D\/3D exploration of large scholarly figure collections, enabling visualization service providers to accelerate trend analysis, technique discovery, and consulting workflows.<\/p>\n","protected":false},"author":0,"featured_media":2610,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[30],"tags":[13,27,17,15],"class_list":["post-386","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-30","tag-machine-learning","tag-systems-design","tag-tech-society","tag-toolservice"],"_links":{"self":[{"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/posts\/386","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/comments?post=386"}],"version-history":[{"count":3,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/posts\/386\/revisions"}],"predecessor-version":[{"id":497,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/posts\/386\/revisions\/497"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/media\/2610"}],"wp:attachment":[{"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/media?parent=386"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/categories?post=386"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/itp.nyu.edu\/thesis\/archive\/2025\/wp-json\/wp\/v2\/tags?post=386"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}