This project is based on the Chinese pre-trained language models bert-base-chinese and xlm-roberta-base, combined with the RNN neural network structure, to build a lightweight Chinese sentiment analysis system. The model is trained on the ChnSentiCorp dataset and supports rapid deployment and use.
The project is suitable for Chinese short-text sentiment classification tasks and can be easily extended to scenarios like comment analysis and user feedback recognition.
The model was trained with 1000 training samples from the ChnSentiCorp dataset for 3 epochs. The accuracy on the full test set is shown below:
| Model Architecture | Accuracy (ACC) |
|---|---|
| BERT (bert-base-chinese) | 88.42% |
| XLM-RoBERTa (xlm-roberta-base) | 87.75% |
| BERT + RNN | 88.92% |
| XLM-RoBERTa + RNN | 88.67% |
- Install dependencies:
pip install -r requirements.txt- Run the demo:
python demo.pyStart the training task using main.py. You can modify training parameters (such as model type, number of epochs, batch size, etc.) in config.py:
python main.py --model bert+rnn --num_epoch 5The trained model will be saved in the ./trained_model directory (the trained model has already been uploaded and can be used directly).
Test the model performance on the validation set:
python test_set.pyUse the trained model to predict sentiment:
python demo.pyYou can input any Chinese sentence, and the model will automatically determine its sentiment (positive/negative).
├── config.py # Hyperparameter configuration
├── main.py # Model training entry point
├── demo.py # Single-sentence sentiment prediction
├── test_set.py # Evaluate the model on the test set
├── model.py # Model definition (including RNN structure)
├── trained_model/ # Directory for saving trained models
├── requirements.txt # Python dependencies
