Skip to content

Latest commit

ย 

History

History
319 lines (253 loc) ยท 8.66 KB

File metadata and controls

319 lines (253 loc) ยท 8.66 KB

Smart Crib Guard

ไธญๆ–‡ | English

License: MIT Python 3.11+

An intelligent baby crib monitoring system based on multimodal large language models and edge-cloud integrated architecture, enabling real-time baby status recognition, risk early warning, and intelligent interaction.

โœจ Key Features

  • ๐ŸŽฏ Baby Behavior Multi-classification๏ผšBased on Qwen2.5-VL-7B LoRA fine-tuning, covering 9 typical behaviors

    • Rolling, climbing crib rails, crying, spitting up, sneezing, yawning, turning, sleeping, quiet state
    • Training Loss converged to 0.0175, excellent model accuracy
  • ๐Ÿ“Š Automated Data Processing Pipeline

    • Video frame extraction with automatic sampling
    • Multiple data augmentation techniques: brightness adjustment, angle transformation, random cropping, flipping
    • Standardized JSON annotation format supporting batch processing
  • ๐Ÿš€ Efficient Inference Deployment

    • Support for single image and batch video inference
    • Real-time classification results with visualization
    • Confusion matrix and classification report generation
  • ๐Ÿ”ง Edge-Cloud Integrated Architecture

    • Local inference combined with cloud-based fine-tuning
    • Flexible configuration management system
    • Support for model hot-updates

๐Ÿš€ Quick Start

Prerequisites

  • Python 3.11+
  • CUDA 12.0+ (GPU recommended)
  • 8GB+ VRAM (16GB+ recommended)

Installation

# Clone the repository
git clone https://github.com/Zsyyxrs/smart-crib-guard.git
cd smart-crib-guard

# Create virtual environment
python -m venv venv
source venv/bin/activate  # Windows: venv\Scripts\activate

# Install dependencies
pip install -r requirements.txt

Quick Inference Examples

Single Image Inference

python scripts/inference/infer_image.py \
  --image_path "path/to/image.jpg" \
  --model_name_or_path "Qwen/Qwen2.5-VL-7B" \
  --adapter_name_or_path "path/to/lora/adapter"

Video Inference

python scripts/inference/infer_video.py \
  --video_path "path/to/video.mp4" \
  --model_name_or_path "Qwen/Qwen2.5-VL-7B" \
  --adapter_name_or_path "path/to/lora/adapter" \
  --output_dir "./output"

Build Dataset

python scripts/data/build_dataset.py \
  --input_dir "path/to/raw/data" \
  --output_json "dataset.json"

๐Ÿ“– Project Structure

smart-crib-guard/
โ”œโ”€โ”€ README.md                    # Chinese documentation (main)
โ”œโ”€โ”€ README_EN.md                 # English documentation
โ”œโ”€โ”€ LICENSE                      # MIT License
โ”œโ”€โ”€ requirements.txt             # Dependencies
โ”œโ”€โ”€ pyproject.toml              # Project configuration
โ”‚
โ”œโ”€โ”€ src/
โ”‚   โ””โ”€โ”€ crib_guard/             # Core package
โ”‚       โ”œโ”€โ”€ __init__.py
โ”‚       โ”œโ”€โ”€ data/               # Data processing module
โ”‚       โ”‚   โ”œโ”€โ”€ augmentation.py # Data augmentation
โ”‚       โ”‚   โ”œโ”€โ”€ dataset.py      # Dataset handling
โ”‚       โ”‚   โ””โ”€โ”€ processing.py   # Data preprocessing
โ”‚       โ”œโ”€โ”€ models/             # Model module
โ”‚       โ”‚   โ””โ”€โ”€ inference.py    # Inference utilities
โ”‚       โ””โ”€โ”€ utils/              # Utility functions
โ”‚           โ””โ”€โ”€ helpers.py      # Helper functions
โ”‚
โ”œโ”€โ”€ scripts/                     # Scripts directory
โ”‚   โ”œโ”€โ”€ train/                  # Training scripts
โ”‚   โ”‚   โ”œโ”€โ”€ train_lora.py       # LoRA fine-tuning
โ”‚   โ”‚   โ””โ”€โ”€ config/             # Training configs
โ”‚   โ”œโ”€โ”€ data/                   # Data processing scripts
โ”‚   โ”‚   โ”œโ”€โ”€ build_dataset.py    # Build dataset
โ”‚   โ”‚   โ”œโ”€โ”€ augment.py          # Data augmentation
โ”‚   โ”‚   โ””โ”€โ”€ merge.py            # Merge datasets
โ”‚   โ””โ”€โ”€ inference/              # Inference scripts
โ”‚       โ”œโ”€โ”€ infer_image.py      # Image inference
โ”‚       โ””โ”€โ”€ infer_video.py      # Video inference
โ”‚
โ”œโ”€โ”€ result/                      # Inference results directory
โ”‚   โ”œโ”€โ”€ .gitkeep
โ”‚   โ””โ”€โ”€ README.md
โ”‚
โ”œโ”€โ”€ output/                      # Model output directory
โ”‚   โ”œโ”€โ”€ .gitkeep
โ”‚   โ””โ”€โ”€ README.md
โ”‚
โ””โ”€โ”€ data/                        # Data directory (local only)
    โ”œโ”€โ”€ .gitkeep
    โ””โ”€โ”€ README.md

๐Ÿ“Š Core Achievements

Model Performance

Metric Value
Base Model Qwen2.5-VL-7B
Fine-tuning Method LoRA (r=64, alpha=16)
Training Loss 0.0175 โ†“
Classification Classes 9 types
Inference Latency ~1.2s/image (GPU)

Behavior Classes

ID Behavior Risk Level Description
1 Rolling Low Baby self-initiated rolling
2 Climbing Rails High Baby attempting to climb crib rails
3 Crying Medium Baby crying state
4 Spitting Up Medium Baby spitting up
5 Sneezing Low Physiological reflex
6 Yawning Low Normal sleep indicator
7 Turning Low Body turning motion
8 Sleeping Low Normal sleep state
9 Quiet Low Awake and quiet

๐Ÿ— Technical Architecture

Data Processing Pipeline

Raw Video
    โ†“
[Frame Extraction] โ†’ Frame Sequence
    โ†“
[Data Augmentation] โ†’ Diverse Dataset
  โ€ข Brightness Adjustment
  โ€ข Angle Transformation
  โ€ข Random Cropping
  โ€ข Flipping
    โ†“
[JSON Annotation] โ†’ Standardized Format
    โ†“
[Dataset Building] โ†’ Training Ready

Model Fine-tuning Pipeline

Qwen2.5-VL-7B (Base Model)
    โ†“
[LoRA Adapter]
    โ†“
[Supervised Fine-tuning] (SFT)
  โ€ข Learning Rate: 1e-4
  โ€ข Batch Size: 32
  โ€ข Optimizer: AdamW
    โ†“
[Convergence Verification] Loss: 0.0175
    โ†“
[Model Merging] โ†’ Production Model

Inference Deployment

Input (Image/Video)
    โ†“
[Preprocessing] โ†’ Qwen2.5-VL Format
    โ†“
[Model Inference] โ†’ Classification Result
    โ†“
[Post-processing] โ†’ Risk Warning
    โ†“
Output (Classification + Visualization)

๐Ÿ“ Usage Examples

1. Build Dataset

from src.crib_guard.data import build_dataset

# Build dataset from raw videos
dataset = build_dataset(
    video_dir="./data/videos",
    annotation_file="./data/annotations.json",
    output_dir="./output/dataset"
)

2. Batch Inference

from src.crib_guard.models import InferenceEngine

engine = InferenceEngine(
    model_name="Qwen/Qwen2.5-VL-7B",
    adapter_path="./output/adapter",
    device="cuda:0"
)

# Inference
results = engine.infer_video("path/to/video.mp4")

# Visualization
engine.visualize_results(results, output_path="./output/results.mp4")

3. Performance Evaluation

python scripts/inference/infer_video.py \
  --video_path "test_video.mp4" \
  --generate_report \
  --save_confusion_matrix

๐Ÿ”ง Configuration Guide

Training Configuration (scripts/train/config)

# Qwen2.5-VL-7B LoRA Fine-tuning Configuration
model_name_or_path: Qwen/Qwen2.5-VL-7B
adapter_name_or_path: null
lora_target: q_proj,v_proj,k_proj,o_proj,gate_proj,up_proj,down_proj
lora_rank: 64
lora_alpha: 16
lora_dropout: 0.1
dataset: baby_behavior
template: qwen_vl
learning_rate: 1e-4
num_train_epochs: 3
batch_size: 32

Inference Configuration

# Inference parameters
config = {
    "model_name": "Qwen/Qwen2.5-VL-7B",
    "adapter_path": "./output/adapter",
    "temperature": 0.3,
    "max_tokens": 128,
    "device": "cuda:0"
}

๐Ÿค Contributing

Contributions are welcome! Please feel free to submit issues and pull requests.

  1. Fork the repository
  2. Create your feature branch (git checkout -b feature/AmazingFeature)
  3. Commit your changes (git commit -m 'Add some AmazingFeature')
  4. Push to the branch (git push origin feature/AmazingFeature)
  5. Open a Pull Request

๐Ÿ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

๐Ÿ‘ค Author

Shangyi Zhu

๐Ÿ“š References

๐ŸŽฏ Future Plans

  • Support for more vision models (LLaVA, GPT-4V)
  • Real-time inference web service
  • Mobile deployment adaptation
  • Large-scale dataset training
  • Behavior prediction and anomaly detection

Last Updated: 2026-05-29

If you have any questions, please open an issue on GitHub Issues.