Skip to content

Repository files navigation

SocialAlign

This repo is the implementation of paper From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media.

Overview

SocialAlign is the first unified framework designed to predicts real-world responses at both micro and macro levels in social contexts. Our framework employs SocialLLM with an articulate Personalized Analyze-Compose LoRA (PAC-LoRA) structure, which deploys specialized expert modules for content analysis and response generation across diverse topics and user profiles, enabling the generation of personalized comments with corresponding sentiments. Experimental results demonstrates that SocialAlign surpasses strong baselines, enhancing public response prediction accuracy in both micro and macro levels while effectively capturing sentiment trends on social media.

Project Structure

  • data_collection includes three subprojects for crawling Weibo data.

    • weibo-ai-search is used for crawling posts from Weibo AI Search URL. As the timestamp of detailed page in Weibo AI Search remains changing, you need to specify a URL of detailed page to be crawled. Run:
    cd data_collection/weibo-ai-search
    python run_spider.py --url "the url you would like to crawl" --output "OUTPUT_FILE_PATH"
  • dataset_construction contains the code for constructing our SocialWeibo dataset.

    We provide some demo cases here.

  • modeling_pac_lora is the implementation of PAC-LoRA, along with the modified Qwen2 model in order to adapt it to our PAC-LoRA architecture and task. pac_lora_layer.py is based on peft 0.12 and another two scripts are based on transformers 4.46.

  • fine-tuning includes the scripts for fine-tuning our SocialLLM and some baseline models.

  • inference includes code to infer baselines and our SocialLLM.

  • utils contains some utility functions used in our project.

We would release all code after notification.

Getting Started

Installation

clone our repo and execute pip install -r requirements.txt to install the requirements needed.

Data Collection

Your can collect hashtagged posts discussing social events through two channals: Weibo Search and Weibo AI Search. weibo-crawler can be utilized to collect user history posts for each unique user appeared in the crawled posts with hashtag.

Dataset Construction

The pipeline of SocialWeibo dataset construction begins with organizing a raw dataset, in which we would remove low-quality user historical posts, clean text noise and then retrieve relevant posts for each user according to the given news content. For example, you may refer to this to construct raw dataset when using Weibo Search as the data source.

After obtaining the raw dataset, we would extract user persona for each user according to the given user historical posts here. Please set your OpenAI API Key before extracting user personas:

export OPENAI_API_KEY="your_openai_api_key"

and then construct SocialWeibo through organize_alphca_dataset.py. Our dataset is in alphca format.

Fine-tuning

Please install our pac-peft and pac-transformers libraries in the environment by changing into the two folders and run pip install -e . respectively.

Then, run the script fine_tuning/fine_tune_pac_lora.py to fine-tuning.

Inference

After obtaining the PAC-LoRA weights, you can perform inference on the test set using infer_socialLLM.py.

Moreover, you do not need to merge weights, as the assembly of multi-analyzing and writing experts is dynamic.

Citation

If you find our work is useful for your research or applications, please kindly cite us:

@inproceedings{10.1145/3746027.3754828,
author = {Zhang, Jinghui and Wan, Kaiyang and Xu, Longwei and Li, Ao and Liu, Zongfang and Chen, Xiuying},
title = {From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media},
year = {2025},
isbn = {9798400720352},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3746027.3754828},
doi = {10.1145/3746027.3754828},
booktitle = {Proceedings of the 33rd ACM International Conference on Multimedia},
pages = {5903–5912},
numpages = {10},
location = {Dublin, Ireland},
series = {MM '25}
}

Acknowledgement

  1. weibo-search and weibo-crawler in data_collection are based on the two projects, respectively:
  2. The implementation of our PAC-LoRA structure is based on Huggingface Transformers and PEFT libraries.

About

[ACM MM 2025] Implementation of paper "From Individuals to Crowds: Dual-Level Public Response Prediction in Social Media"

Topics

Resources

Stars

8 stars

Watchers

1 watching

Forks

Packages

Contributors

Languages