llm-inference

🔮 SuperDuperDB: Bring AI to your database! Build, deploy and manage any AI application directly with your existing data infrastructure, without moving your data. Including streaming inference, scalable model training and vector search.

Updated May 5, 2024
Python

neuralmagic / deepsparse

Star

Sparsity-aware deep learning inference runtime for CPUs

nlp performance computer-vision inference machinelearning pruning object-detection pretrained-models quantization cpus onnx sparsification llm-inference deepsparse

Updated May 2, 2024
Python

InternLM / lmdeploy

Star

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

llama cuda-kernels deepspeed llm fastertransformer llm-inference turbomind internlm llama2 codellama llama3

Updated May 5, 2024
Python

databricks / dbrx

Star

Code examples and resources for DBRX, a large language model developed by Databricks

databricks llm generative-ai gen-ai llm-training llm-inference mosaic-ai

Updated May 1, 2024
Python

intel / intel-extension-for-transformers

Star

⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡

cpu retrieval gpu chatbot xeon rag habana large-language-model chatpdf llm-inference 4-bits speculative-decoding llm-cpu streamingllm neurips2023 intel-optimized-llamacpp neural-chat gaudi2 neural-chat-7b

Updated Apr 30, 2024
Python

liltom-eth / llama2-webui

Star

Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.

llm llm-inference llama2 llama-2

Updated Mar 22, 2024
Jupyter Notebook

FasterDecoding / Medusa

Star

Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

llm llm-inference

Updated Apr 18, 2024
Jupyter Notebook

microsoft / aici

Star

AICI: Prompts as (Wasm) Programs

rust ai wasm inference transformer language-model model-serving wasmtime llm llmops llm-serving llm-inference llm-framework

Updated May 4, 2024
Rust

yomorun / yomo

Sponsor

Star

🦖 Stateful Serverless Framework for building Geo-distributed Edge AI Infra

serverless realtime webassembly stream-processing low-latency quic edge-computing geo-distributed geodistributedsystems yomo edge-ai distributed-cloud llm-inference function-calling

Updated Apr 30, 2024
Go

predibase / lorax

Star

Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

transformers pytorch llama gpt lora model-serving fine-tuning llm llmops llm-serving llm-inference

Updated May 4, 2024
Python

NVIDIA / GenerativeAIExamples

Star

Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

microservice gpu-acceleration nemo tensorrt rag triton-inference-server large-language-models llm llm-inference retrieval-augmented-generation

Updated Apr 22, 2024
Python

DefTruth / Awesome-LLM-Inference

Star

📖A curated list of Awesome LLM Inference Paper with codes, TensorRT-LLM, vLLM, streaming-llm, AWQ, SmoothQuant, WINT8/4, Continuous Batching, FlashAttention, PagedAttention etc.

sora awq llm llms vllm llm-inference awesome-llm flash-attention flash-attention-2 tensorrt-llm paged-attention streaming-llm streamingllm flash-decoding inferflow

Updated May 2, 2024

Improve this page

Add a description, image, and links to the llm-inference topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the llm-inference topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

llm-inference

Here are 369 public repositories matching this topic...

nomic-ai / gpt4all

microsoft / autogen

bentoml / OpenLLM

mistralai / mistral-src

SJTU-IPADS / PowerInfer

Lightning-AI / litgpt

liguodongiot / llm-action

openvinotoolkit / openvino

SuperDuperDB / superduperdb

neuralmagic / deepsparse

InternLM / lmdeploy

databricks / dbrx

intel / intel-extension-for-transformers

liltom-eth / llama2-webui

FasterDecoding / Medusa

microsoft / aici

yomorun / yomo

predibase / lorax

NVIDIA / GenerativeAIExamples

DefTruth / Awesome-LLM-Inference

Improve this page

Add this topic to your repo