Choosing Local Model for Vision

Final Quiz

Connecting to LMS... Progress: in progress

Assessment

1. What distinguishes a vision-language model from a text-only language model?
2. What should be defined first when selecting a local vision model?
3. Which resource most directly constrains whether a GPU can hold a local vision model and its active workload?
4. When can a smaller model be the better operational choice?
5. What does model quantization primarily change?
6. Which statement about Q8 and Q4 quantization is most accurate?
7. Why should image-token or visual-context settings be included in a benchmark configuration?
8. How can thinking or reasoning mode affect perception-heavy tasks?
9. What is a visual hallucination?
10. Why should the same model configuration be run more than once during evaluation?
11. What makes a local vision benchmark representative?
12. Which metric captures whether a model successfully returns usable output across the test set?
13. What should be recorded to make a local model benchmark reproducible?
14. Which deployment practice best supports responsible use of local vision models?
15. Which statement best summarizes local vision model selection?