LocalLLaMA

1

30

Beginner questions thread (self.localllama)

submitted 1 year ago by noneabove1182 to c/localllama

23 comments fedilink

Trying something new, going to pin this thread as a place for beginners to ask what may or may not be stupid questions, to encourage both the asking and answering.

Depending on activity level I'll either make a new one once in awhile or I'll just leave this one up forever to be a place to learn and ask.

When asking a question, try to make it clear what your current knowledge level is and where you may have gaps, should help people provide more useful concise answers!

2

76

Free Open-Source AI LLM Guide (lemmy.world)

submitted 2 years ago by [email protected] to c/localllama

4 comments fedilink

cross-posted from: https://lemmy.world/post/2219010

Hello everyone!

We have officially hit 1,000 subscribers! How exciting!! Thank you for being a member of [email protected]. Whether you're a casual passerby, a hobby technologist, or an up-and-coming AI developer - I sincerely appreciate your interest and support in a future that is free and open for all.

It can be hard to keep up with the rapid developments in AI, so I have decided to pin this at the top of our community to be a frequently updated LLM-specific resource hub and model index for all of your adventures in FOSAI.

The ultimate goal of this guide is to become a gateway resource for anyone looking to get into free open-source AI (particularly text-based large language models). I will be doing a similar guide for image-based diffusion models soon!

In the meantime, I hope you find what you're looking for! Let me know in the comments if there is something I missed so that I can add it to the guide for everyone else to see.

Getting Started With Free Open-Source AI

Have no idea where to begin with AI / LLMs? Try starting with our Lemmy Crash Course for Free Open-Source AI.

When you're ready to explore more resources see our FOSAI Nexus - a hub for all of the major FOSS & FOSAI on the cutting/bleeding edges of technology.

If you're looking to jump right in, I recommend downloading oobabooga's text-generation-webui and installing one of the LLMs from TheBloke below.

Try both GGML and GPTQ variants to see which model type performs to your preference. See the hardware table to get a better idea on which parameter size you might be able to run (3B, 7B, 13B, 30B, 70B).

8-bit System Requirements

Model VRAM Used Minimum Total VRAM Card Examples RAM/Swap to Load*

LLaMA-7B 9.2GB 10GB 3060 12GB, 3080 10GB 24 GB

LLaMA-13B 16.3GB 20GB 3090, 3090 Ti, 4090 32 GB

LLaMA-30B 36GB 40GB A6000 48GB, A100 40GB 64 GB

LLaMA-65B 74GB 80GB A100 80GB 128 GB

4-bit System Requirements

Model Minimum Total VRAM Card Examples RAM/Swap to Load*

LLaMA-7B 6GB GTX 1660, 2060, AMD 5700 XT, RTX 3050, 3060 6 GB

LLaMA-13B 10GB AMD 6900 XT, RTX 2060 12GB, 3060 12GB, 3080, A2000 12 GB

LLaMA-30B 20GB RTX 3080 20GB, A4500, A5000, 3090, 4090, 6000, Tesla V100 32 GB

LLaMA-65B 40GB A100 40GB, 2x3090, 2x4090, A40, RTX A6000, 8000 64 GB

*System RAM (not VRAM), is utilized to initially load a model. You can use swap space if you do not have enough RAM to support your LLM.

When in doubt, try starting with 3B or 7B models and work your way up to 13B+.

FOSAI Resources

Fediverse / FOSAI

The Internet is Healing

FOSAI Welcome Message

FOSAI Crash Course

FOSAI Nexus Resource Hub

LLM Leaderboards

HF Open LLM Leaderboard

LMSYS Chatbot Arena

LLM Search Tools

LLM Explorer

Open LLMs

Model	VRAM Used	Minimum Total VRAM	Card Examples	RAM/Swap to Load*
LLaMA-7B	9.2GB	10GB	3060 12GB, 3080 10GB	24 GB
LLaMA-13B	16.3GB	20GB	3090, 3090 Ti, 4090	32 GB
LLaMA-30B	36GB	40GB	A6000 48GB, A100 40GB	64 GB
LLaMA-65B	74GB	80GB	A100 80GB	128 GB

Model	Minimum Total VRAM	Card Examples	RAM/Swap to Load*
LLaMA-7B	6GB	GTX 1660, 2060, AMD 5700 XT, RTX 3050, 3060	6 GB
LLaMA-13B	10GB	AMD 6900 XT, RTX 2060 12GB, 3060 12GB, 3080, A2000	12 GB
LLaMA-30B	20GB	RTX 3080 20GB, A4500, A5000, 3090, 4090, 6000, Tesla V100	32 GB
LLaMA-65B	40GB	A100 40GB, 2x3090, 2x4090, A40, RTX A6000, 8000	64 GB

Large Language Model Hub

Download Models

oobabooga

text-generation-webui - a big community favorite gradio web UI by oobabooga designed for running almost any free open-source and large language models downloaded off of HuggingFace which can be (but not limited to) models like LLaMA, llama.cpp, GPT-J, Pythia, OPT, and many others. Its goal is to become the AUTOMATIC1111/stable-diffusion-webui of text generation. It is highly compatible with many formats.

Exllama

A standalone Python/C++/CUDA implementation of Llama for use with 4-bit GPTQ weights, designed to be fast and memory-efficient on modern GPUs.

gpt4all

Open-source assistant-style large language models that run locally on your CPU. GPT4All is an ecosystem to train and deploy powerful and customized large language models that run locally on consumer-grade processors.

TavernAI

The original branch of software SillyTavern was forked from. This chat interface offers very similar functionalities but has less cross-client compatibilities with other chat and API interfaces (compared to SillyTavern).

SillyTavern

Developer-friendly, Multi-API (KoboldAI/CPP, Horde, NovelAI, Ooba, OpenAI+proxies, Poe, WindowAI(Claude!)), Horde SD, System TTS, WorldInfo (lorebooks), customizable UI, auto-translate, and more prompt options than you'd ever want or need. Optional Extras server for more SD/TTS options + ChromaDB/Summarize. Based on a fork of TavernAI 1.2.8

Koboldcpp

A self contained distributable from Concedo that exposes llama.cpp function bindings, allowing it to be used via a simulated Kobold API endpoint. What does it mean? You get llama.cpp with a fancy UI, persistent stories, editing tools, save formats, memory, world info, author's note, characters, scenarios and everything Kobold and Kobold Lite have to offer. In a tiny package around 20 MB in size, excluding model weights.

KoboldAI-Client

This is a browser-based front-end for AI-assisted writing with multiple local & remote AI models. It offers the standard array of tools, including Memory, Author's Note, World Info, Save & Load, adjustable AI settings, formatting options, and the ability to import existing AI Dungeon adventures. You can also turn on Adventure mode and play the game like AI Dungeon Unleashed.

h2oGPT

h2oGPT is a large language model (LLM) fine-tuning framework and chatbot UI with document(s) question-answer capabilities. Documents help to ground LLMs against hallucinations by providing them context relevant to the instruction. h2oGPT is fully permissive Apache V2 open-source project for 100% private and secure use of LLMs and document embeddings for document question-answer.

Models

The Bloke

The Bloke is a developer who frequently releases quantized (GPTQ) and optimized (GGML) open-source, user-friendly versions of AI Large Language Models (LLMs).

These conversions of popular models can be configured and installed on personal (or professional) hardware, bringing bleeding-edge AI to the comfort of your home.

Support TheBloke here.

https://ko-fi.com/TheBlokeAI

70B

Llama-2-70B-chat-GPTQ

Llama-2-70B-Chat-GGML

Llama-2-70B-GPTQ

Llama-2-70B-GGML

llama-2-70b-Guanaco-QLoRA-GPTQ

30B

30B-Epsilon-GPTQ

13B

Llama-2-13B-chat-GPTQ

Llama-2-13B-chat-GGML

Llama-2-13B-GPTQ

Llama-2-13B-GGML

llama-2-13B-German-Assistant-v2-GPTQ

llama-2-13B-German-Assistant-v2-GGML

13B-Ouroboros-GGML

13B-Ouroboros-GPTQ

13B-BlueMethod-GGML

13B-BlueMethod-GPTQ

llama-2-13B-Guanaco-QLoRA-GGML

llama-2-13B-Guanaco-QLoRA-GPTQ

Dolphin-Llama-13B-GGML

Dolphin-Llama-13B-GPTQ

MythoLogic-13B-GGML

MythoBoros-13B-GPTQ

WizardLM-13B-V1.2-GPTQ

WizardLM-13B-V1.2-GGML

OpenAssistant-Llama2-13B-Orca-8K-3319-GGML

7B

Llama-2-7B-GPTQ

Llama-2-7B-GGML

Llama-2-7b-Chat-GPTQ

LLongMA-2-7B-GPTQ

llama-2-7B-Guanaco-QLoRA-GPTQ

llama-2-7B-Guanaco-QLoRA-GGML

llama2_7b_chat_uncensored-GPTQ

llama2_7b_chat_uncensored-GGML

More Models

Any of KoboldAI's Models

Luna-AI-Llama2-Uncensored-GPTQ

Nous-Hermes-Llama2-GGML

Nous-Hermes-Llama2-GPTQ

FreeWilly2-GPTQ

GL, HF!

Are you an LLM Developer? Looking for a shoutout or project showcase? Send me a message and I'd be more than happy to share your work and support links with the community.

If you haven't already, consider subscribing to the free open-source AI community at [email protected] where I will do my best to make sure you have access to free open-source artificial intelligence on the bleeding edge.

Thank you for reading!

3

16

Trying out old GPUs with Vulkan (discuss.tchncs.de)

submitted 7 hours ago by [email protected] to c/localllama

11 comments fedilink

Yesterday I got bored and decided to try out my old GPUs with Vulkan. I had an HD 5830, GTX 460 and GTX 770 4Gb laying around so I figured "Why not".

Long story short - Vulkan didn't recognize them, hell, Linux didn't even recognize them. They didn't show up in nvtop, nvidia-smi or anything. I didn't think to check dmesg.

Honestly, I thought the 770 would work; it hasn't been in legacy status that long. It might work with an older Nvidia driver version (I'm on 550 now) but I'm not messing with that stuff just because I'm bored.

So for now the oldest GPUs I can get running are a Ryzen 5700G APU and 1080ti. Both Vega and Pascal came out in early 2017 according to Wikipedia. Those people disappointed that their RX 500 and RX 5000 don't work in Ollama should give Llama.cpp Vulkan a shot. Kobold has a Vulkan option too.

The 5700G works fine alongside Nvidia GPUs in Vulkan. The performance is what you'd expect from an APU, but at least it works. Now I'm tempted to buy a 7600 XT just to see how it does.

Has anyone else out there tried Vulkan?

4

39

AMD denies rumors of Radeon RX 9070 XT with 32GB memory (videocardz.com)

submitted 4 days ago by [email protected] to c/localllama

9 comments fedilink

Well, it was nice ... having hope, I mean. That was a good feeling.

5

10

Models not loading into RAM (lemmy.ml)

submitted 3 days ago by [email protected] to c/localllama

9 comments fedilink

I didn't expect a 8B-F16 model with 16GB on disk could be run in my laptop with only 16GB of RAM and integrated GPU, It was painfuly slow, like 0.3 t/s, but it ran. Then I learnt that you can effectively run a model from your storage without loading into memory and checked that it was exactly the case, the memory usage kept constant at around 20% with and without running the model. The problem is that gpt4all-chat is running all the models greater than 1.5B in this way, and the difference is huge as the 1.5b model runs at 20 t/s. Even a distilled 6.7B_Q8 model with roughly 7GB on disk that has plenty of room (12GB RAM free) didn't move the memory usage and it was also very slow (3 tokens/sec). I'm pretty new to this field so I'm probably missing something basic, but I just followed the instrucctions for downloading it and compile it.

6

57

AMD reportedly working on gaming Radeon RX 9070 XT GPU with 32GB memory (videocardz.com)

submitted 6 days ago* (last edited 6 days ago) by [email protected] to c/localllama

18 comments fedilink

One might question why an RX 9070 card would need so much memory, but increased capacity can serve purposes beyond gaming, such as Large Language Model (LLM) support for AI workloads. Additionally, it’s worth noting that RX 9070 cards will use 20 Gbps memory, much slower than the RTX 50 series, which features 28-30 Gbps GDDR7 variants. So, while capacity may increase, bandwidth likely won’t.

7

Recommend models for GTX 1660 Super (6GB) (lemmy.sdf.org)

submitted 5 days ago by [email protected] to c/localllama

6 comments fedilink

I have an GTX 1660 Super (6 GB)

Right now I have ollama with:

deepseek-r1:8b
qwen2.5-coder:7b

Do you recommend any other local models to play with my GPU?

8

12

The Anthropic Economic Index - an initiative aimed at understanding AI's effects on labor markets and the economy over time. (www.anthropic.com)

submitted 1 week ago by [email protected] to c/localllama

3 comments fedilink

9

4

AI Action Summit in Paris (www.youtube.com)

submitted 1 week ago by [email protected] to c/localllama

3 comments fedilink

Closing session, speech by Modi, JD Vance, Ursula von der Leyen

10

6

French President Emmanuel Macron announces €100 billion investments in AI (www.france24.com)

submitted 1 week ago by [email protected] to c/localllama

0 comments fedilink

11

8

"Flash Answers" Cerebras brings instant inference to Mistral Le Chat (cerebras.ai)

submitted 1 week ago by [email protected] to c/localllama

1 comments fedilink

Sorry I keep posting about Mistral but if you check: https://chat.mistral.ai/chat

I duno how they do it but some of these answers are lightning fast:

Fast inference dramatically improves the user experience for chat and code generation – two of the most popular use-cases today. In the example above, Mistral Le Chat completes a coding prompt instantly while other popular AI assistants take up to 50 seconds to finish.

For this initial release, Cerebras will focus on serving text-based queries for the Mistral Large 2 model. When using Cerebras Inference, Le Chat will display a “Flash Answer ⚡” icon on the bottom left of the chat interface.

12

7

Hibiki by kyutai, a simultaneous speech-to-speech translation model, currently supporting FR to EN (aussie.zone)

submitted 1 week ago* (last edited 1 week ago) by [email protected] to c/localllama

1 comments fedilink

Example of it working in action: https://streamable.com/ueh3sj

Paper: https://arxiv.org/abs/2502.03382

Samples: https://hf.co/spaces/kyutai/hibiki-samples

Inference code: https://github.com/kyutai-labs/hibiki

Models: https://huggingface.co/kyutai

From kyutai on X: Meet Hibiki, our simultaneous speech-to-speech translation model, currently supporting FR to EN.

Hibiki produces spoken and text translations of the input speech in real-time, while preserving the speaker’s voice and optimally adapting its pace based on the semantic content of the source speech.

Based on objective and human evaluations, Hibiki outperforms previous systems for quality, naturalness and speaker similarity and approaches human interpreters.

https://x.com/kyutai_labs/status/1887495488997404732

Neil Zeghidour on X: https://x.com/neilzegh/status/1887498102455869775

13

16

DeepSeek gives Europe's tech firms a chance to catch up in global AI race (www.reuters.com)

submitted 2 weeks ago by [email protected] to c/localllama

6 comments fedilink

14

34

How to run LLaMA (and other LLMs) on Android. (lemmy.dbzer0.com)

submitted 2 weeks ago* (last edited 2 weeks ago) by [email protected] to c/localllama

17 comments fedilink

Hello, everyone! I wanted to share my experience of successfully running LLaMA on an Android device. The model that performed the best for me was llama3.2:1b on a mid-range phone with around 8 GB of RAM. I was also able to get it up and running on a lower-end phone with 4 GB RAM. However, I also tested several other models that worked quite well, including qwen2.5:0.5b , qwen2.5:1.5b , qwen2.5:3b , smallthinker , tinyllama , deepseek-r1:1.5b , and gemma2:2b. I hope this helps anyone looking to experiment with these models on mobile devices!

Step 1: Install Termux

Download and install Termux from the Google Play Store or F-Droid

Step 2: Set Up proot-distro and Install Debian

Open Termux and update the package list:
```
pkg update && pkg upgrade
```
Install proot-distro
```
pkg install proot-distro
```
Install Debian using proot-distro:
```
proot-distro install debian
```
Log in to the Debian environment:
```
proot-distro login debian
```
You will need to log-in every time you want to run Ollama. You will need to repeat this step and all the steps below every time you want to run a model (excluding step 3 and the first half of step 4).

Step 3: Install Dependencies

Update the package list in Debian:
```
apt update && apt upgrade
```
Install curl:
```
apt install curl
```

Step 4: Install Ollama

Run the following command to download and install Ollama:
```
curl -fsSL https://ollama.com/install.sh | sh
```
Start the Ollama server:
```
ollama serve &
```
After you run this command, do ctrl + c and the server will continue to run in the background.

Step 5: Download and run the Llama3.2:1B Model

Use the following command to download the Llama3.2:1B model:
```
ollama run llama3.2:1b
```
This step fetches and runs the lightweight 1-billion-parameter version of the Llama 3.2 model .

Running LLaMA and other similar models on Android devices is definitely achievable, even with mid-range hardware. The performance varies depending on the model size and your device's specifications, but with some experimentation, you can find a setup that works well for your needs. I’ll make sure to keep this post updated if there are any new developments or additional tips that could help improve the experience. If you have any questions or suggestions, feel free to share them below!

– llama

15

What is a good model that runs on 6GB Vram? (discuss.online)

submitted 2 weeks ago by [email protected] to c/localllama

10 comments fedilink

Should be good at conversations and creative, it'll be for worldbuilding

Best if uncensored as I prefer that over it kicking in when I least want it

I'm fine with those roleplaying models as long as they can actually give me ideas and talk to be logically

16

12

Has anyone applied tree of thought prompting to r1 yet? (programming.dev)

submitted 2 weeks ago by [email protected] to c/localllama

9 comments fedilink

Generate 5 thoughts, prune 3, branch, repeat. I think that’s what o1 pro and o3 do

17

24

Mistral Small 3 (24B) released (mistral.ai)

submitted 2 weeks ago by [email protected] to c/localllama

1 comments fedilink

18

15

Did DeepSeek R1 just pop nvidias bubble? (www.youtube.com)

submitted 3 weeks ago by [email protected] to c/localllama

8 comments fedilink

Changed title because no need for youtube clickbait here

19

28

Why llms are suprisingly good at math, and what it means to process language. (lemmy.world)

submitted 3 weeks ago* (last edited 3 weeks ago) by [email protected] to c/localllama

20 comments fedilink

Someone asked about how llms can be so good at math operations. My response comment kind of turned into a five paragraph essay as they tend to do sometimes. Thought I would offer it here and add some reference. Maybe spark some discussion?

What do language models do?

LLMs are trained to recognize, process, and construct patterns of language data into high dimensional manifold plots.

Meaning its job is to structure and compartmentalize the patterns of language into a map where each word and its particular meaning live as pairs of points on a geometric surface. Its point is placed near closely related points in space connected by related concepts or properties of the word.

You can explore such a map for vision models here!

Then they use that map to statistically navigate through the sea of ways words can be associated into sentences to find coherent paths.

What does language really mean?

Language data isnt just words and syntax, its underlying abstract concepts, context, and how humans choose to compartmentalize or represent universal ideas given our subjective reference point.

Language data extends to everything humans can construct thoughts about including mathematics, philosophy, science storytelling, music theory, programming, ect.

Language is universal because its a fundimental way we construct and organize concepts. The first important cognative milestone for babies is the association of concepts to words and constructing sentences with them.

Even the universe speaks its own language. Physical reality and logical abstractions speak the same underlying universal patterns hidden in formalized truths and dynamical operation. Information and matter are two sides to a coin, their structure is intrinsicallty connected.

Math and conceptual vectors

Math is a symbolic representation of combinatoric logic. Logic is generally a formalized language used to represent ideas related to truth as well as how truth can be built on through axioms.

Numbers and math is cleanly structured and formalized patterns of language data. Its riggerously described and its axioms well defined. So its relatively easy to train a model to recognize and internalize patterns inherent to basic arithmetic and linear algebra and how they manipulate or process the data points representing numbers.

You can imagine the llms data manifold having a section for math and logic processing. The concept of one lives somewhere as a point of data on the manifold. By moving a point representing the concept of one along a vector dimension that represents the process of 'addition by one' to find the data point representing two.

Not a calculator though

However an llm can never be a true calculator due to the statistical nature of the tokenizer. It always has a chance of giving the wrong answer. In the infinite multitude of tokens it can pick any number of wrong numbers. We can get the statistical chance of failure down though.

Its an interesting how llms can still give accurate answers for artithmatic despite having no in built calculation function. Through training alone they are learning how to apply simple arithmetic.

hidden structures of information

There are hidden or intrinsic patterns to most structures of information. Usually you can find the fractal hyperstructures the patterns are geometrically baked into in higher dimensions once you go plotting out their phase space/ holomorphic parameter maps. We can kind of visualize these fractals with vision model activation parameter maps. Welch labs on yt has a great video about it.

Modern language models have so many parameters with so many dimensions the manifold expands into its impossible to visualize. So they are basically mystery black boxes that somehow understand these crazy fractal structures of complex information and navigate the topological manifolds language data creates.

conclusion

This is my understanding of how llms do their thing. I hope you enjoyed reading! Secretly I just wanted to show you the cool chart :)

20

25

Thoughts on new deepseek R1 distill models (lemmy.world)

submitted 3 weeks ago* (last edited 3 weeks ago) by [email protected] to c/localllama

7 comments fedilink

Ive been playing around with the deepseek R1 distills. Qwen 14b and 32b specifically.

So far its very cool to see models really going after this current CoT meta by mimicing internal thinking monologues. Seeing a model go "but wait..." "Hold on, let me check again..." "Aha! So.." Kind of makes it feel more natural in its eventual conclusions.

I don't like how it can get caught in looping thought processes and im not sure how much all the extra tokens spent really go towards a "better" answer/solution.

What really needs to be ironed out is the reading comprehension seems to be lower th average as it misses small details in tricky questions and makes assumptions about what youre trying to ask like wanting a recipe for coconut oil cookies but only seeing coconut and giving a coconut cookie recipe with regular butter.

Its exciting to see models operate in a kind of a new way.

21

13

unsure on how to quantize model (feddit.it)

submitted 1 month ago by [email protected] to c/localllama

5 comments fedilink

I was experimenting with oobabooga trying to run this model but due to it's size it wasn't going to fit in ram, so i tried to quantize it using llama.cpp, and that worked, but due to the gguf format it was only running on the cpu. searching for ways to quantize the model while keeping it in safetensors returned nothing; so is there any way to do that?

I'm sorry if this is a stupid question, i still know almost nothing of this field

22

12

How much gpu do i need to run a 90b model (lemm.ee)

submitted 1 month ago by [email protected] to c/localllama

16 comments fedilink

Do i need industry grade gpu's or can i scrape by getring decent tps with a consumer level gpu.

23

7

Nvidia Digits AI Supercomputer just announced (lemmy.world)

submitted 1 month ago by [email protected] to c/localllama

0 comments fedilink