Project brief
A simple view of what needs to be delivered.
Build the underlying system that allows an intelligence engineer to transform large volumes of public fan conversation into structured, evidence-backed intelligence. Create an initial production-quality pipeline capable of collecting, processing, classifying and synthesizing fan conversations from authorized/public data sources. The initial version concentrate primarily on YouTube videos and comments. The architecture should later allow the platform to add other approved sources. The system should identify signals including: Sentiment Emotional reaction Fan requests Praise Criticism Questions Repeated themes Cultural references Collaboration requests Song requests Tour requests Merch requests Release anticipation Nostalgia Fan theories Memes Frequently quoted lyrics or moments Emerging topics Sudden changes in conversation Unusual engagement patterns Recurring fan pain points Content fans want more of The vendor should not simply generate generic LLM summaries. The system should preserve the underlying evidence that generated every important insight. Required Data Architecture At minimum capture: Artist Video type Publishing date Comment date Comment likes Parent/reply relationship Language Engagement signals Topic classifications Sentiment classifications Entity references Relevant extracted phrases Time-series signal Supporting evidence for generated insights Deliverables Fully documented ingestion pipeline. Normalized data model. Processing and cleaning pipeline. Duplicate/spam detection. Language detection. Topic extraction system. Sentiment/emotion classification. Acceptance Criteria Pipeline works end-to-end. New artists can be added without engineering changes. Data processing is repeatable. Insights can be queried through an API/database. Supporting evidence is retained. System has been demonstrated on real-world artist datasets. Documentation is sufficient for another engineer to operate the system. Suggested Duration 4–6 weeks.
