OpenVINO
/

mixtral-8x7b-instruct-v0.1-int4-ov

@@ -27,10 +27,10 @@ For more information on quantization, check the [OpenVINO model optimization gui
 The provided OpenVINO™ IR model is compatible with:
-* OpenVINO version 2024.0.0 and higher
 * Optimum Intel 1.16.0 and higher
-## Running Model Inference
 1. Install packages required for using [Optimum Intel](https://huggingface.co/docs/optimum/intel/index) integration with the OpenVINO backend:
@@ -49,20 +49,47 @@ tokenizer = AutoTokenizer.from_pretrained(model_id)
 model = OVModelForCausalLM.from_pretrained(model_id)
-messages = [
-    {"role": "user", "content": "What is your favourite condiment?"},
-    {"role": "assistant", "content": "Well, I'm quite partial to a good squeeze of fresh lemon juice. It adds just the right amount of zesty flavour to whatever I'm cooking up in the kitchen!"},
-    {"role": "user", "content": "Do you have mayonnaise recipes?"}
-]
-inputs = tokenizer.apply_chat_template(messages, return_tensors="pt")
-outputs = model.generate(inputs, max_new_tokens=20)
-print(tokenizer.decode(outputs[0], skip_special_tokens=True))
 ```
 For more examples and possible optimizations, refer to the [OpenVINO Large Language Model Inference Guide](https://docs.openvino.ai/2024/learn-openvino/llm_inference_guide.html).
 ## Limitations
 Check the original model card for [limitations](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1#limitations).

 The provided OpenVINO™ IR model is compatible with:
+* OpenVINO version 2024.2.0 and higher
 * Optimum Intel 1.16.0 and higher
+## Running Model Inference with [Optimum Intel](https://huggingface.co/docs/optimum/intel/index)
 1. Install packages required for using [Optimum Intel](https://huggingface.co/docs/optimum/intel/index) integration with the OpenVINO backend:
 model = OVModelForCausalLM.from_pretrained(model_id)
+inputs = tokenizer("What is OpenVINO?", return_tensors="pt")
+outputs = model.generate(**inputs, max_length=200)
+text = tokenizer.batch_decode(outputs)[0]
+print(text)
 ```
 For more examples and possible optimizations, refer to the [OpenVINO Large Language Model Inference Guide](https://docs.openvino.ai/2024/learn-openvino/llm_inference_guide.html).
+## Running Model Inference with [OpenVINO GenAI](https://github.com/openvinotoolkit/openvino.genai)
+1. Install packages required for using OpenVINO GenAI.
+```
+pip install openvino-genai huggingface_hub
+```
+2. Download model from HuggingFace Hub
+```
+import huggingface_hub as hf_hub
+model_id = "OpenVINO/mixtral-8x7b-instruct-v0.1-int4-ov"
+model_path = "mixtral-8x7b-instruct-v0.1-int4-ov"
+hf_hub.snapshot_download(model_id, local_dir=model_path)
+```
+3. Run model inference:
+```
+import openvino_genai as ov_genai
+device = "CPU"
+pipe = ov_genai.LLMPipeline(model_path, device)
+print(pipe.generate("What is OpenVINO?"))
+```
+More GenAI usage examples can be found in OpenVINO GenAI library [docs](https://github.com/openvinotoolkit/openvino.genai/blob/master/src/README.md) and [samples](https://github.com/openvinotoolkit/openvino.genai?tab=readme-ov-file#openvino-genai-samples)
 ## Limitations
 Check the original model card for [limitations](https://huggingface.co/mistralai/Mixtral-8x7B-Instruct-v0.1#limitations).