ollama: init package (langchain-ai#23615)

Co-authored-by: Erick Friis <erick@langchain.dev>
olgamurraft · Aug 16, 2024 · 1933e22 · 1933e22
1 parent 068c5f4
commit 1933e22
Showing 28 changed files with 8,726 additions and 510 deletions.
diff --git a/docs/docs/integrations/chat/ollama.ipynb b/docs/docs/integrations/chat/ollama.ipynb
diff --git a/docs/docs/integrations/llms/ollama.ipynb b/docs/docs/integrations/llms/ollama.ipynb
@@ -1,227 +1,125 @@
 {
  "cells": [
+  {
+   "cell_type": "raw",
+   "id": "67db2992",
+   "metadata": {},
+   "source": [
+    "---\n",
+    "sidebar_label: Ollama\n",
+    "---"
+   ]
+  },
   {
    "cell_type": "markdown",
+   "id": "9597802c",
    "metadata": {},
    "source": [
-    "# Ollama\n",
+    "# OllamaLLM\n",
     "\n",
     ":::caution\n",
     "You are currently on a page documenting the use of Ollama models as [text completion models](/docs/concepts/#llms). Many popular Ollama models are [chat completion models](/docs/concepts/#chat-models).\n",
     "\n",
     "You may be looking for [this page instead](/docs/integrations/chat/ollama/).\n",
     ":::\n",
     "\n",
-    "[Ollama](https://ollama.ai/) allows you to run open-source large language models, such as Llama 2, locally.\n",
-    "\n",
-    "Ollama bundles model weights, configuration, and data into a single package, defined by a Modelfile. \n",
-    "\n",
-    "It optimizes setup and configuration details, including GPU usage.\n",
-    "\n",
-    "For a complete list of supported models and model variants, see the [Ollama model library](https://github.com/ollama/ollama#model-library).\n",
+    "This page goes over how to use LangChain to interact with `Ollama` models.\n",
     "\n",
+    "## Installation"
+   ]
+  },
+  {
+   "cell_type": "code",
+   "execution_count": null,
+   "id": "59c710c4",
+   "metadata": {},
+   "outputs": [],
+   "source": [
+    "# install package\n",
+    "%pip install -U langchain-ollama"
+   ]
+  },
+  {
+   "cell_type": "markdown",
+   "id": "0ee90032",
+   "metadata": {},
+   "source": [
     "## Setup\n",
     "\n",
-    "First, follow [these instructions](https://github.com/ollama/ollama) to set up and run a local Ollama instance:\n",
+    "First, follow [these instructions](https://github.com/jmorganca/ollama) to set up and run a local Ollama instance:\n",
     "\n",
     "* [Download](https://ollama.ai/download) and install Ollama onto the available supported platforms (including Windows Subsystem for Linux)\n",
     "* Fetch available LLM model via `ollama pull <name-of-model>`\n",
-    "    * View a list of available models via the [model library](https://ollama.ai/library) and pull to use locally with the command `ollama pull llama3`\n",
+    "    * View a list of available models via the [model library](https://ollama.ai/library)\n",
+    "    * e.g., `ollama pull llama3`\n",
     "* This will download the default tagged version of the model. Typically, the default points to the latest, smallest sized-parameter model.\n",
     "\n",
     "> On Mac, the models will be download to `~/.ollama/models`\n",
     "> \n",
     "> On Linux (or WSL), the models will be stored at `/usr/share/ollama/.ollama/models`\n",
     "\n",
     "* Specify the exact version of the model of interest as such `ollama pull vicuna:13b-v1.5-16k-q4_0` (View the [various tags for the `Vicuna`](https://ollama.ai/library/vicuna/tags) model in this instance)\n",
-    "* To view all pulled models on your local instance, use `ollama list`\n",
+    "* To view all pulled models, use `ollama list`\n",
     "* To chat directly with a model from the command line, use `ollama run <name-of-model>`\n",
-    "* View the [Ollama documentation](https://github.com/ollama/ollama) for more commands. \n",
-    "* Run `ollama help` in the terminal to see available commands too.\n",
-    "\n",
-    "## Usage\n",
-    "\n",
-    "You can see a full list of supported parameters on the [API reference page](https://api.python.langchain.com/en/latest/llms/langchain_community.llms.ollama.Ollama.html).\n",
+    "* View the [Ollama documentation](https://github.com/jmorganca/ollama) for more commands. Run `ollama help` in the terminal to see available commands too.\n",
     "\n",
-    "If you are using a LLaMA `chat` model (e.g., `ollama pull llama3`) then you can use the `ChatOllama` [interface](https://python.langchain.com/v0.2/docs/integrations/chat/ollama/).\n",
-    "\n",
-    "This includes [special tokens](https://ollama.com/library/llama3) for system message and user input.\n",
-    "\n",
-    "## Interacting with Models \n",
-    "\n",
-    "Here are a few ways to interact with pulled local models\n",
-    "\n",
-    "#### In the terminal:\n",
-    "\n",
-    "* All of your local models are automatically served on `localhost:11434`\n",
-    "* Run `ollama run <name-of-model>` to start interacting via the command line directly\n",
-    "\n",
-    "#### Via the API\n",
-    "\n",
-    "Send an `application/json` request to the API endpoint of Ollama to interact.\n",
-    "\n",
-    "```bash\n",
-    "curl http://localhost:11434/api/generate -d '{\n",
-    "  \"model\": \"llama3\",\n",
-    "  \"prompt\":\"Why is the sky blue?\"\n",
-    "}'\n",
-    "```\n",
-    "\n",
-    "See the Ollama [API documentation](https://github.com/ollama/ollama/blob/main/docs/api.md) for all endpoints.\n",
-    "\n",
-    "#### via LangChain\n",
-    "\n",
-    "See a typical basic example of using [Ollama chat model](https://python.langchain.com/v0.2/docs/integrations/chat/ollama/) in your LangChain application."
+    "## Usage"
    ]
   },
   {
    "cell_type": "code",
-   "execution_count": null,
-   "metadata": {},
-   "outputs": [],
-   "source": [
-    "!pip install langchain-community"
-   ]
-  },
-  {
-   "cell_type": "code",
-   "execution_count": 1,
-   "metadata": {},
+   "execution_count": 4,
+   "id": "035dea0f",
+   "metadata": {
+    "tags": []
+   },
    "outputs": [
     {
      "data": {
       "text/plain": [
-       "\"Here's one:\\n\\nWhy don't scientists trust atoms?\\n\\nBecause they make up everything!\\n\\nHope that made you smile! Do you want to hear another one?\""
+       "'A great start!\\n\\nLangChain is a type of AI model that uses language processing techniques to generate human-like text based on input prompts or chains of reasoning. In other words, it can have a conversation with humans, understanding the context and responding accordingly.\\n\\nHere\\'s a possible breakdown:\\n\\n* \"Lang\" likely refers to its focus on natural language processing (NLP) and linguistic analysis.\\n* \"Chain\" suggests that LangChain is designed to generate text in response to a series of connected ideas or prompts, rather than simply generating random text.\\n\\nSo, what do you think LangChain\\'s capabilities might be?'"
       ]
      },
-     "execution_count": 1,
+     "execution_count": 4,
      "metadata": {},
      "output_type": "execute_result"
     }
    ],
    "source": [
-    "from langchain_community.llms import Ollama\n",
+    "from langchain_core.prompts import ChatPromptTemplate\n",
+    "from langchain_ollama.llms import OllamaLLM\n",
     "\n",
-    "llm = Ollama(\n",
-    "    model=\"llama3\"\n",
-    ")  # assuming you have Ollama installed and have llama3 model pulled with `ollama pull llama3 `\n",
+    "template = \"\"\"Question: {question}\n",
     "\n",
-    "llm.invoke(\"Tell me a joke\")"
-   ]
-  },
-  {
-   "cell_type": "markdown",
-   "metadata": {},
-   "source": [
-    "To stream tokens, use the `.stream(...)` method:"
-   ]
-  },
-  {
-   "cell_type": "code",
-   "execution_count": 4,
-   "metadata": {},
-   "outputs": [
-    {
-     "name": "stdout",
-     "output_type": "stream",
-     "text": [
-      "\n",
-      "\n",
-      "S\n",
-      "ure\n",
-      ",\n",
-      " here\n",
-      "'\n",
-      "s\n",
-      " one\n",
-      ":\n",
-      "\n",
-      "\n",
-      "\n",
-      "\n",
-      "Why\n",
-      " don\n",
-      "'\n",
-      "t\n",
-      " scient\n",
-      "ists\n",
-      " trust\n",
-      " atoms\n",
-      "?\n",
-      "\n",
-      "\n",
-      "B\n",
-      "ecause\n",
-      " they\n",
-      " make\n",
-      " up\n",
-      " everything\n",
-      "!\n",
-      "\n",
-      "\n",
-      "\n",
-      "\n",
-      "I\n",
-      " hope\n",
-      " you\n",
-      " found\n",
-      " that\n",
-      " am\n",
-      "using\n",
-      "!\n",
-      " Do\n",
-      " you\n",
-      " want\n",
-      " to\n",
-      " hear\n",
-      " another\n",
-      " one\n",
-      "?\n",
-      "\n"
-     ]
-    }
-   ],
-   "source": [
-    "query = \"Tell me a joke\"\n",
+    "Answer: Let's think step by step.\"\"\"\n",
     "\n",
-    "for chunks in llm.stream(query):\n",
-    "    print(chunks)"
-   ]
-  },
-  {
-   "cell_type": "markdown",
-   "metadata": {},
-   "source": [
-    "To learn more about the LangChain Expressive Language and the available methods on an LLM, see the [LCEL Interface](/docs/concepts#interface)"
+    "prompt = ChatPromptTemplate.from_template(template)\n",
+    "\n",
+    "model = OllamaLLM(model=\"llama3\")\n",
+    "\n",
+    "chain = prompt | model\n",
+    "\n",
+    "chain.invoke({\"question\": \"What is LangChain?\"})"
    ]
   },
   {
    "cell_type": "markdown",
+   "id": "e2d85456",
    "metadata": {},
    "source": [
     "## Multi-modal\n",
     "\n",
-    "Ollama has support for multi-modal LLMs, such as [bakllava](https://ollama.ai/library/bakllava) and [llava](https://ollama.ai/library/llava).\n",
+    "Ollama has support for multi-modal LLMs, such as [bakllava](https://ollama.com/library/bakllava) and [llava](https://ollama.com/library/llava).\n",
     "\n",
-    "`ollama pull bakllava`\n",
+    "    ollama pull bakllava\n",
     "\n",
     "Be sure to update Ollama so that you have the most recent version to support multi-modal."
    ]
   },
-  {
-   "cell_type": "code",
-   "execution_count": 5,
-   "metadata": {},
-   "outputs": [],
-   "source": [
-    "from langchain_community.llms import Ollama\n",
-    "\n",
-    "bakllava = Ollama(model=\"bakllava\")"
-   ]
-  },
   {
    "cell_type": "code",
    "execution_count": 2,
+   "id": "4043e202",
    "metadata": {},
    "outputs": [
     {
@@ -279,7 +177,8 @@
   },
   {
    "cell_type": "code",
-   "execution_count": 8,
+   "execution_count": 4,
+   "id": "79aaf863",
    "metadata": {},
    "outputs": [
     {
@@ -288,38 +187,24 @@
        "'90%'"
       ]
      },
-     "execution_count": 8,
+     "execution_count": 4,
      "metadata": {},
      "output_type": "execute_result"
     }
    ],
    "source": [
-    "llm_with_image_context = bakllava.bind(images=[image_b64])\n",
-    "llm_with_image_context.invoke(\"What is the dollar based gross retention rate:\")"
-   ]
-  },
-  {
-   "cell_type": "markdown",
-   "metadata": {},
-   "source": [
-    "## Concurrency Features\n",
+    "from langchain_ollama import OllamaLLM\n",
     "\n",
-    "Ollama supports concurrency inference for a single model, and or loading multiple models simulatenously (at least [version 0.1.33](https://github.com/ollama/ollama/releases)).\n",
+    "llm = OllamaLLM(model=\"bakllava\")\n",
     "\n",
-    "Start the Ollama server with:\n",
-    "\n",
-    "* `OLLAMA_NUM_PARALLEL`: Handle multiple requests simultaneously for a single model\n",
-    "* `OLLAMA_MAX_LOADED_MODELS`: Load multiple models simultaneously\n",
-    "\n",
-    "Example: `OLLAMA_NUM_PARALLEL=4 OLLAMA_MAX_LOADED_MODELS=4 ollama serve`\n",
-    "\n",
-    "Learn more about configuring Ollama server in [the official guide](https://github.com/ollama/ollama/blob/main/docs/faq.md#how-do-i-configure-ollama-server)."
+    "llm_with_image_context = llm.bind(images=[image_b64])\n",
+    "llm_with_image_context.invoke(\"What is the dollar based gross retention rate:\")"
    ]
   }
  ],
  "metadata": {
   "kernelspec": {
-   "display_name": "Python 3 (ipykernel)",
+   "display_name": "Python 3.11.1 64-bit",
    "language": "python",
    "name": "python3"
   },
@@ -333,9 +218,14 @@
    "name": "python",
    "nbconvert_exporter": "python",
    "pygments_lexer": "ipython3",
-   "version": "3.11.8"
+   "version": "3.12.3"
+  },
+  "vscode": {
+   "interpreter": {
+    "hash": "e971737741ff4ec9aff7dc6155a1060a59a8a6d52c757dbbe66bf8ee389494b1"
+   }
   }
  },
  "nbformat": 4,
- "nbformat_minor": 4
+ "nbformat_minor": 5
 }