Conversation
|
hi @Yavaren, thanks for the pr. i think this warrants a careful approach, do you have a visual of what this looks like? and i think mlx-vlm downstream requires datasets>=2.19.1 as well for lora |
|
Hello @Lazarus-931! I attached all of the visuals. The adaptor selector is available through the model configuration panel or models page. I had a talk with Prince today regarding the looks and usability and got positive feedback. I am really looking forward to hear what you think!
|
|
I think this is nice, lgtm! merge conflicts btw |
6b2b456 to
4d785ac
Compare
|
@Lazarus-931 fixed the conflicts |
|
new conflicts! |
| from typing import Any, Mapping, Sequence | ||
|
|
||
|
|
||
| LORA_SUFFIXES = (".lora_a", ".lora_b") |
There was a problem hiding this comment.
Should this live in mlx-vlm instead, or is there a reason it has to be an overlay?
There was a problem hiding this comment.
I agree we should add mlx-vlm long-term. Right now standard PEFT adapters are packaged differently, so their settings and weights need to be converted before mlx-vlm can use them. It also loads weights with strict=False so unmatched keys can be silently skipped.
There was a problem hiding this comment.
But the conversion is just remapping keys, right? I think this should just live in mlx-vlm since it's useful there too, and generally we're trying to keep the python overlay as minimal as possible here so we can consolidate effort for fixing bugs and maintenance of models in the server instead of split across two different surface areas. FWIW I don't think this should be a major lift, there's already some amount of LoRA support in mlx-vlm today that we can extend.
There was a problem hiding this comment.
Yes, you are right. I had a call with Prince, he showed me something. Thanks a lot for pointing this out!
There was a problem hiding this comment.
Lucas, do you have any specific plans for extending LoRA support in mlx-vlm?
There was a problem hiding this comment.
@Yavaren I didn't, but feel free to ping me if you want to chat about it or have a specific request.
|
@Yavaren Looks cool! I had a couple of questions/comments inline but overall this seems good. |
# Conflicts: # PythonDistribution/Requirements/mlx-vlm-server-macos-arm64.txt # Sources/Nativ/Features/Chat/ChatConfigurationView.swift
…-lora-adapters # Conflicts: # Makefile # Nativ.xcodeproj/project.pbxproj # Nativ.xcodeproj/project.xcworkspace/xcshareddata/swiftpm/Package.resolved # PythonDistribution/Overlay/nativ_server.py # PythonDistribution/Scripts/build_mlx_vlm_server.py # README.md # Sources/Nativ/ControlPanelView.swift # Sources/Nativ/Features/Chat/ChatConfigurationView.swift # Sources/Nativ/Features/Models/HuggingFaceHub.swift # Sources/Nativ/Features/Models/ModelsView.swift # Sources/Nativ/NativModel.swift # project.yml




What changed
This PR adds LoRA adapter support for language models.
Adapters can now be found and downloaded from Hugging Face directly from the model configuration panel. Installed adapters can be selected, activated, paused during download, or deleted.
On the server side, adapters are validated before loading, including their configuration and compatibility with the selected base model. GGUF repositories are excluded.
I also added clearer download progress and error handling.
Current scope
This supports adapters downloaded from Hugging Face only. Importing adapters from arbitrary local directories is intentionally not included in this PR.