ไธญๆ | English
An intelligent baby crib monitoring system based on multimodal large language models and edge-cloud integrated architecture, enabling real-time baby status recognition, risk early warning, and intelligent interaction.
-
๐ฏ Baby Behavior Multi-classification๏ผBased on Qwen2.5-VL-7B LoRA fine-tuning, covering 9 typical behaviors
- Rolling, climbing crib rails, crying, spitting up, sneezing, yawning, turning, sleeping, quiet state
- Training Loss converged to 0.0175, excellent model accuracy
-
๐ Automated Data Processing Pipeline
- Video frame extraction with automatic sampling
- Multiple data augmentation techniques: brightness adjustment, angle transformation, random cropping, flipping
- Standardized JSON annotation format supporting batch processing
-
๐ Efficient Inference Deployment
- Support for single image and batch video inference
- Real-time classification results with visualization
- Confusion matrix and classification report generation
-
๐ง Edge-Cloud Integrated Architecture
- Local inference combined with cloud-based fine-tuning
- Flexible configuration management system
- Support for model hot-updates
- Python 3.11+
- CUDA 12.0+ (GPU recommended)
- 8GB+ VRAM (16GB+ recommended)
# Clone the repository
git clone https://github.com/Zsyyxrs/smart-crib-guard.git
cd smart-crib-guard
# Create virtual environment
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtpython scripts/inference/infer_image.py \
--image_path "path/to/image.jpg" \
--model_name_or_path "Qwen/Qwen2.5-VL-7B" \
--adapter_name_or_path "path/to/lora/adapter"python scripts/inference/infer_video.py \
--video_path "path/to/video.mp4" \
--model_name_or_path "Qwen/Qwen2.5-VL-7B" \
--adapter_name_or_path "path/to/lora/adapter" \
--output_dir "./output"python scripts/data/build_dataset.py \
--input_dir "path/to/raw/data" \
--output_json "dataset.json"smart-crib-guard/
โโโ README.md # Chinese documentation (main)
โโโ README_EN.md # English documentation
โโโ LICENSE # MIT License
โโโ requirements.txt # Dependencies
โโโ pyproject.toml # Project configuration
โ
โโโ src/
โ โโโ crib_guard/ # Core package
โ โโโ __init__.py
โ โโโ data/ # Data processing module
โ โ โโโ augmentation.py # Data augmentation
โ โ โโโ dataset.py # Dataset handling
โ โ โโโ processing.py # Data preprocessing
โ โโโ models/ # Model module
โ โ โโโ inference.py # Inference utilities
โ โโโ utils/ # Utility functions
โ โโโ helpers.py # Helper functions
โ
โโโ scripts/ # Scripts directory
โ โโโ train/ # Training scripts
โ โ โโโ train_lora.py # LoRA fine-tuning
โ โ โโโ config/ # Training configs
โ โโโ data/ # Data processing scripts
โ โ โโโ build_dataset.py # Build dataset
โ โ โโโ augment.py # Data augmentation
โ โ โโโ merge.py # Merge datasets
โ โโโ inference/ # Inference scripts
โ โโโ infer_image.py # Image inference
โ โโโ infer_video.py # Video inference
โ
โโโ result/ # Inference results directory
โ โโโ .gitkeep
โ โโโ README.md
โ
โโโ output/ # Model output directory
โ โโโ .gitkeep
โ โโโ README.md
โ
โโโ data/ # Data directory (local only)
โโโ .gitkeep
โโโ README.md
| Metric | Value |
|---|---|
| Base Model | Qwen2.5-VL-7B |
| Fine-tuning Method | LoRA (r=64, alpha=16) |
| Training Loss | 0.0175 โ |
| Classification Classes | 9 types |
| Inference Latency | ~1.2s/image (GPU) |
| ID | Behavior | Risk Level | Description |
|---|---|---|---|
| 1 | Rolling | Low | Baby self-initiated rolling |
| 2 | Climbing Rails | High | Baby attempting to climb crib rails |
| 3 | Crying | Medium | Baby crying state |
| 4 | Spitting Up | Medium | Baby spitting up |
| 5 | Sneezing | Low | Physiological reflex |
| 6 | Yawning | Low | Normal sleep indicator |
| 7 | Turning | Low | Body turning motion |
| 8 | Sleeping | Low | Normal sleep state |
| 9 | Quiet | Low | Awake and quiet |
Raw Video
โ
[Frame Extraction] โ Frame Sequence
โ
[Data Augmentation] โ Diverse Dataset
โข Brightness Adjustment
โข Angle Transformation
โข Random Cropping
โข Flipping
โ
[JSON Annotation] โ Standardized Format
โ
[Dataset Building] โ Training Ready
Qwen2.5-VL-7B (Base Model)
โ
[LoRA Adapter]
โ
[Supervised Fine-tuning] (SFT)
โข Learning Rate: 1e-4
โข Batch Size: 32
โข Optimizer: AdamW
โ
[Convergence Verification] Loss: 0.0175
โ
[Model Merging] โ Production Model
Input (Image/Video)
โ
[Preprocessing] โ Qwen2.5-VL Format
โ
[Model Inference] โ Classification Result
โ
[Post-processing] โ Risk Warning
โ
Output (Classification + Visualization)
from src.crib_guard.data import build_dataset
# Build dataset from raw videos
dataset = build_dataset(
video_dir="./data/videos",
annotation_file="./data/annotations.json",
output_dir="./output/dataset"
)from src.crib_guard.models import InferenceEngine
engine = InferenceEngine(
model_name="Qwen/Qwen2.5-VL-7B",
adapter_path="./output/adapter",
device="cuda:0"
)
# Inference
results = engine.infer_video("path/to/video.mp4")
# Visualization
engine.visualize_results(results, output_path="./output/results.mp4")python scripts/inference/infer_video.py \
--video_path "test_video.mp4" \
--generate_report \
--save_confusion_matrix# Qwen2.5-VL-7B LoRA Fine-tuning Configuration
model_name_or_path: Qwen/Qwen2.5-VL-7B
adapter_name_or_path: null
lora_target: q_proj,v_proj,k_proj,o_proj,gate_proj,up_proj,down_proj
lora_rank: 64
lora_alpha: 16
lora_dropout: 0.1
dataset: baby_behavior
template: qwen_vl
learning_rate: 1e-4
num_train_epochs: 3
batch_size: 32# Inference parameters
config = {
"model_name": "Qwen/Qwen2.5-VL-7B",
"adapter_path": "./output/adapter",
"temperature": 0.3,
"max_tokens": 128,
"device": "cuda:0"
}Contributions are welcome! Please feel free to submit issues and pull requests.
- Fork the repository
- Create your feature branch (
git checkout -b feature/AmazingFeature) - Commit your changes (
git commit -m 'Add some AmazingFeature') - Push to the branch (
git push origin feature/AmazingFeature) - Open a Pull Request
This project is licensed under the MIT License - see the LICENSE file for details.
Shangyi Zhu
- GitHub: @Zsyyxrs
- Email: y5bcgb98fr@privaterelay.appleid.com
- Support for more vision models (LLaVA, GPT-4V)
- Real-time inference web service
- Mobile deployment adaptation
- Large-scale dataset training
- Behavior prediction and anomaly detection
Last Updated: 2026-05-29
If you have any questions, please open an issue on GitHub Issues.