VStar is a workflow that utilizes visual contextual information to process high-definition images. It utilizes hierarchical image structures and adaptive thresholding to efficiently identify objects relevant to user queries.
This example demonstrates how to use the OMAgent framework for visual search and analysis tasks. The example code can be found in the "examples/VStar" directory.
cd examples/VStarThis example implements a comprehensive VStar workflow that consists of the following components:
-
VStar Input
- Handles user input containing text queries and image uploads
- Processes multi-modal inputs to prepare for visual analysis
-
Vstar Workflow
- Determine the elements needed to answer the question
- Performing confidence-guided searches to localize visual elements
- Optimize search efficiency using adaptive thresholding
- Python 3.11+
- Required packages installed (see requirements.txt)
- Access to a multimodal LLM (e.g., LLaVA, GPT-4V) or compatible endpoint
- Redis server running locally or remotely (for pro mode)
- Conductor server running locally or remotely (for pro mode)
The container.yaml file manages dependencies and settings for different components of the system. To set up your configuration:
-
Generate the
container.yamlfile:python compile_container.py
This will create a
container.yamlfile with default settings underexamples/VStar. -
Configure your multimodal LLM settings in
configs/llms/*.yml:- Set your model endpoint through environment variables or by directly modifying the yml file
export custom_vstar_endpoint="your_vstar_endpoint"
- Configure other model settings like temperature as needed.
-
Update settings in the generated
container.yaml:- Modify Redis connection settings (for pro mode):
- Set the host, port, and credentials for your Redis instance.
- Configure both
redis_stream_clientandredis_stm_clientsections.
- Update the Conductor server URL under the conductor_config section (for pro mode).
- Adjust any other component settings as needed.
- Modify Redis connection settings (for pro mode):
Run the VStar example:
For terminal/CLI usage:
python run_cli.pyYou can run the VStar workflow in pro mode or lite mode by changing the OMAGENT_MODE environment variable. The default mode is pro, which uses the conductor and Redis server. The lite mode will run the workflow in the current Python process without external services.
For pro mode:
export OMAGENT_MODE="pro"
python run_cli.pyFor lite mode:
export OMAGENT_MODE="lite"
python run_cli.pyVStar uses a hierarchical approach to image analysis:
- The image is first processed to extract relevant features and prepare for analysis.
- Visual cues are generated from the user's query to guide the search.
- A confidence-guided search algorithm traverses the image data to locate visual elements.
- Adaptive thresholding ensures high-quality results while optimizing computation.
- Found elements are synthesized into a comprehensive answer.
This approach enables precise localization of visual elements while maintaining computational efficiency.
If you encounter issues:
- Verify your multimodal LLM endpoint is accessible and working.
- For pro mode, confirm Redis is running and accessible.
- Ensure all dependencies are installed correctly.
- Check for sufficient GPU resources if using local model deployment.
- Review logs for any error messages.
- Open an issue on GitHub if you can't find a solution; we will do our best to help you out!
Since vstar does not yet support deployment by vllm, for example, we need to deploy locally.
First of all, go to V*'s code repository and download the source code
, then copy the python file OmAgent/examples/Vstar/docs/files/vstar_api.py for deploying the api to the vstar source folder, and change the model path to your download seal models. Finally run uvicorn vstar_api:app --host 0.0.0.0 --port 8000 to start the service, and then export custom_vstar_endpoint=http://localhost:8000/.
