Kaia Gao

Qianwen (Kaia)Gao

Understanding people. Evaluating AI. Building useful tools.

Product Analytics · AI Evaluation · User Research · Applied AI

I study how people use technology, evaluate how AI behaves, and build tools that help people make better decisions. My work connects computational social science at UC Berkeley with product experimentation, research accepted to COLM, and hands-on AI projects.

CHASE · Accepted to COLM 2026 →

About Me

People, information, and the systems that connect them

My background in Communication at Zhejiang University led me to ask how information shapes people's choices. Through product analytics at Didi and RedNote, I began studying those questions with behavioral data and experiments. I earned my master's in Computational Social Science at UC Berkeley in May 2026, bringing together research methods, statistics, and computing.

At Wrodium, I investigated how AI systems shape information quality, co-authoring CHASE, accepted to COLM 2026. Alongside research, I build AI applications and have worked on projects involving student absenteeism and youth wellness. I want my work to make technology more trustworthy, help people access useful information, and support decisions that improve their lives.

Product Analytics & Experimentation

Using behavioral data and experiments to inform product decisions. At Didi and RedNote, I worked on user research, pricing analysis, audience segmentation, and A/B testing.

Explore my experience →

AI Evaluation & Integrity

Investigating how ranking incentives and content freshness affect information quality. My work at Wrodium includes CHASE, accepted to COLM 2026, and the FreshRAG benchmark project.

Read my publication →

Quantitative User Research

Combining surveys, behavioral analysis, and statistical methods to understand people's needs. My experience spans transportation, youth wellness, and education.

See my research experience →

Applied AI Engineering

Putting research into practice through software projects, including PickMem for LLM memory curation and Consentful Civic Lens for event consent and storytelling.

Explore my projects →

Publications

COLM 2026

CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target

Qianwen Gao, Zichang Su, Yiwen Hou, Arlen Kumar, Leanid Palkhouski

Accepted to the Conference on Language Modeling (COLM), 2026

A simulation framework examining how repeated optimization for LLM rankings changes content ecosystems. Across six domains, alignment between rankings and independently assessed content quality declines as creators adapt to ranking signals.

Read paper on arXiv

Selected Projects

Evaluating AI, understanding behavior, and building tools people can use

🧠

Applied AI · LLM Memory

PickMem - A local-first memory-curation layer for LLMs

Personal Project | July 2026

Built an AI-powered memory management application that helps users capture, organize, and retrieve information using semantic search and intelligent categorization. Developed a full-stack architecture with secure authentication, optimized backend APIs and database queries for fast retrieval, and delivered a responsive interface for seamless knowledge management.

ReactTypeScriptNode.jsExpressMongoDBMCP
🧪

AI Evaluation · Experimental Design

FreshRAG: Causal Benchmark for RAG Freshness & Hallucination

Research Project | May 2026

Designed FreshRAG, a large-scale causal benchmark (50K+ QA pairs) to measure how content freshness reduces hallucination in retrieval-augmented generation (RAG) systems. Built a temporal-gradient dataset from multi-year knowledge snapshots and constructed controlled retrieval scenarios to isolate mechanisms including knowledge conflict resolution, temporal grounding, and parametric override. Implemented counterfactual evaluation protocols and mechanism-level effect decomposition, enabling regression-based and experimental estimation of freshness treatment effects across models and domains.

PythonRAGCausal InferenceExperimental DesignNLPLLM Evaluation
🎓

Predictive Analytics · Education

Predictive Absenteeism & Early-Warning Signal Analysis

Capstone Project | ONGB & Wizearly | Feb 2026 – May 2026

Collaborated with Oakland Natives Give Back (ONGB) and Wizearly to analyze chronic absenteeism trends by synthesizing national datasets (NCES, Census) with thousands of granular OUSD student records. Engineered a Predictive Feature Library using the IPIR framework to identify behavioral and academic risk signals, validating national patterns against local data. Developed a dual-scale landscape report and interactive dashboard to provide data-driven intervention strategies for school district leadership.

PythonSQLPandasScikit-LearnTableauStatistical Modeling
🎭

Applied AI · Consent

Consentful Civic Lens – Event Organizer

CalHacks Project | Oct 2025

Built a full-stack web app with Next.js, Supabase, and PostgreSQL for event consent management and storytelling. Integrated Claude API and LiveKit for AI-generated highlight summaries, and developed a recommendation system to personalize future event suggestions based on user interests and location.

Next.jsSupabasePostgreSQLClaude APILiveKit
👗

Product Analytics · NLP

Consumer Sentiment & Brand Insights from Amazon Fashion Reviews

Course Project | Oct 2025 – Nov 2025

Analyzed 2.5M Amazon Fashion reviews to extract customer sentiment and brand perception using NLP techniques (VADER, BERT embeddings, topic modeling). Built regression and clustering models to identify key drivers of satisfaction and differentiate brand positioning. Visualized sentiment and keyword trends across categories through an interactive Streamlit dashboard, providing actionable insights for marketing and product strategy.

PythonVADERBERTStreamlitScikit-learn
📈

Behavioral Analysis · Econometrics

Finfluencers Impact on trading behavior

Course Project | Nov 2025 – Dec 2025

Investigated the causal impact of "finfluencer" (financial influencer) sentiment on stock trading liquidity using a balanced panel dataset of five major tech stocks (AAPL, AMZN, FB, NVDA, TSLA) from 2020 to 2022. Constructed a Panel OLS regression model with Entity Fixed Effects and clustered standard errors to control for unobserved heterogeneity and serial correlation. Identified that market volatility (VIX) and negative retail sentiment ("fear") are the primary drivers of trading volume, with the final model explaining 41% of day-to-day variance in trading activity.

Panel OLS RegressionFixed Effects ModelingHypothesis TestingEconometricsStatistical Analysis
🏡

Data Analysis · Housing

California Housing Market Affordability Analysis

Course Project | Nov 2025 – Dec 2025

Investigated the "Gravity of Affordability" in California housing markets by synthesizing construction permit data (HUD), sales volume (Redfin), and demographic trends (NIH) from 1980–2022. Calculated Price-to-Income Ratios (PIR) to quantify affordability gaps across key counties like San Francisco and Riverside, revealing a decoupling of local incomes from housing costs. Visualized supply inelasticity and migration pressures using R (ggplot2) to demonstrate how low affordability drives population shifts despite stagnant construction responsiveness.

Rggplot2dplyrHexData Visualization

Experience & Leadership

Product decisions, research questions, and the people behind them

AI Research Intern

Wrodium

Berkeley, CA · Dec 2025 – July 2026

  • •COLM 2026 Research – Co-authored CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target, accepted to the Conference on Language Modeling (COLM) 2026. Studied how repeated optimization for LLM rankings reshapes content ecosystems across six domains, finding declining alignment between rankings and independently assessed content quality.
  • •Causal Benchmark Development – Led development of a research framework to quantify how content freshness affects LLM hallucination through three causal mechanisms (Knowledge Conflict, Temporal Grounding, Parametric Override).
  • •Temporal QA Dataset Construction – Built QA dataset using Myers diff for factual change detection; designed factorial experiments with logistic regression decomposition to isolate mechanism effects across 6 domains and multiple LLMs.
  • •Content Pipeline Automation – Engineered a multi-agent workflow using Make.com and LLM APIs to automate technical blog generation on Generative Engine Optimization (GEO), synthesizing retrieval-augmented generation (RAG) research into educational content.

Machine Learning Engineer Intern

Oakland Natives Give Back & Wizearly

Oakland, CA · Feb 2025 – May 2026

  • •Built a leakage-safe early-warning pipeline over 256,865 student-year records representing 36,695 OUSD students across seven academic years; engineered 63 attendance, academic, demographic, school, and neighborhood features.
  • •Trained and evaluated eight class-imbalance-aware models using five-fold cross-validation and a temporal holdout; tuned XGBoost achieved 98.9% recall, 0.843 AUC-ROC, and 0.734 PR-AUC.
  • •Tuned the intervention threshold to retain 94.5% recall while reducing the flagged population from 85.3% to 66.6%, cutting false positives by approximately 22% relative to the default threshold.
  • •Developed a predictive-feature library, landscape report, and interactive dashboard to support school-district intervention planning.

Strategy & Data Analyst Intern

APPA Health

Berkeley, CA · Sept – Dec 2025

  • •Market Opportunity Analysis – Analyzed the educational funding landscape to identify and evaluate a pipeline of potential funding opportunities supporting youth wellness.
  • •Impact Measurement & Reporting – Established a KPI framework to measure SEL program effectiveness. Analyzed pre- and post-program survey data to quantify impact on student engagement, providing key insights for program iteration and reporting to funding partners.

Marketing Analytics Intern

RedNote

Shanghai, China · Aug 2024 – Jan 2025

  • •Audience Segmentation – Queried and analyzed behavioral and demographic user data using SQL in Hive on a large-scale data warehouse to create 35 pet industry audience segments, contributing to ¥1.83M (~$250K) in ad revenue and improved ad targeting accuracy within the first month.
  • •KPI Automation – Developed and automated marketing KPI dashboards using Python, SQL, and RedBI (BI tool comparable to Power BI) to track campaign performance, user engagement, and retention metrics. Presented findings and strategic recommendations to over 740 clients and internal stakeholders.
  • •Marketing Strategy – Designed and analyzed A/B tests to optimize ad targeting strategies and creatives. Integrated CRM data to conduct deep-dive analyses on marketing performance, providing insights that improved marketing efficiency and ROI.

Product & User Analytics Intern

Didi

Hangzhou, China · Mar – Jun 2024

  • •Pricing Analytics – Conducted multivariate regression and causal inference analyses on supply-demand patterns and user price elasticity to inform dynamic pricing strategies, leading to a 2% revenue lift.
  • •User Research – Designed and distributed user surveys to identify pain points in the "hourly driver" service; combined findings with SQL-based behavioral analysis to uncover actionable product insights, driving a 3% reduction in complaints and measurable improvement in driver-passenger experience.

Content Operation Intern

Huace Film & TV

Hangzhou, China · Jun – Sept 2023

  • •Content Engagement Analysis – Queried and analyzed 10,000+ follower records using SQL and Python to identify audience attributes and content preferences; created user clusters that informed strategy adjustments, boosting page views by 15.1%.
  • •A/B Testing – Conducted A/B tests to refine video strategy; produced and distributed 300+ YouTube clips, leveraging insights to drive engagement from 620K+ global followers.

President

ZJU Lingyun Musical Club

Hangzhou, China · Sept 2021 – May 2024

  • •Managed club operations across 8 departments with 150+ members; led the annual musical theatre production, drawing 6,000+ audience members.
  • •Produced an original musical commemorating the 40th anniversary of Chu Kochen Honors College, overseeing recruitment, script development, budgeting, and cross-team coordination.

Methods & Tools

The methods and tools I use across research, analysis, and software projects

Product Analytics & Experimentation

SQL & Python
A/B Testing
Causal Inference & Regression
Segmentation & Product Metrics
Tableau & Streamlit

AI Evaluation & Integrity

LLM Evaluation
Benchmark & Dataset Design
RAG & Content Freshness
Factorial Experiments
NLP & Text Analysis

Quantitative User Research

Survey Design
Behavioral Data Analysis
Statistical Modeling in R & Python
Impact Measurement
Research Communication

Applied AI Engineering

Python & TypeScript
React & Next.js
Node.js & REST APIs
PostgreSQL & Supabase
LLM APIs & MCP

Education

Master of Computational Social Science

University of California, Berkeley

Berkeley, CA · Jun 2025 – May 2026

  • •GPA: 3.87/4.00
  • •Relevant Coursework: Advanced Computing, Machine Learning, Advanced Applied Statistics, Data Visualization, Deep Learning for Visual Data (DeCal)

Bachelor of Arts, Communication

Zhejiang University (ZJU)

Hangzhou, China · Sept 2021 – Jun 2025

  • •GPA: 3.95/4.00
  • •Relevant Coursework: Big Data Analytics, Advanced Mathematics, Probability and Mathematical Statistics, Python Programming, Introduction to Research Methodology in Social Sciences

Resume

A one-page overview of my experience in product analytics, AI evaluation, user research, and applied AI.

KG

Qianwen (Kaia) Gao

Computational Social Science · Product Analytics & AI · Berkeley, CA

Get In Touch

I'm exploring early-career opportunities in product analytics, AI evaluation and integrity, quantitative user research, and applied AI engineering. I'd love to connect with teams building useful, trustworthy technology.

Contact Information