Skip to content
Navigation Menu
Sign in
Appearance settings
Platform
AI CODE CREATION
GitHub Copilot
Write better code with AI
GitHub Copilot app
Direct agents from issue to merge
MCP Registry
Integrate external tools
DEVELOPER WORKFLOWS
Actions
Automate any workflow
Codespaces
Instant dev environments
Issues
Plan and track work
Code Review
Manage code changes
Code Quality
Enforce quality at merge
APPLICATION SECURITY
GitHub Advanced Security
Find and fix vulnerabilities
Code security
Secure your code as you build
Secret protection
Stop leaks before they start
EXPLORE
Why GitHub
Documentation
Blog
Changelog
Marketplace
View all features
Solutions
BY COMPANY SIZE
Enterprises
Small and medium teams
Startups
Nonprofits
BY USE CASE
App Modernization
DevSecOps
DevOps
CI/CD
View all use cases
BY INDUSTRY
Healthcare
Financial services
Manufacturing
Government
View all industries
View all solutions
Resources
EXPLORE BY TOPIC
AI
Software Development
DevOps
Security
View all topics
EXPLORE BY TYPE
Customer stories
Events & webinars
Ebooks & reports
Business insights
GitHub Skills
SUPPORT & SERVICES
Documentation
Customer support
Community forum
Trust center
Partners
View all resources
Open Source
COMMUNITY
GitHub Sponsors
Fund open source developers
PROGRAMS
Security Lab
Maintainer Community
Accelerator
GitHub Stars
Archive Program
REPOSITORIES
Topics
Trending
Collections
Enterprise
ENTERPRISE SOLUTIONS
Enterprise platform
AI-powered developer platform
AVAILABLE ADD-ONS
GitHub Advanced Security
Enterprise-grade security features
Copilot for Business
Enterprise-grade AI features
Premium Support
Enterprise-grade 24/7 support
Pricing
Type
/
to search
Sign in
Sign up
Appearance settings
You signed in with another tab or window.
Reload
to refresh your session.
You signed out in another tab or window.
Reload
to refresh your session.
You switched accounts on another tab or window.
Reload
to refresh your session.
Dismiss alert
{{ message }}
gary149
/
llama-agent
Public
forked from
ggml-org/llama.cpp
Notifications
You must be signed in to change notification settings
Fork
29
Star
279
Code
Issues
0
Pull requests
1
Actions
Projects
Security and quality
0
Insights
Additional navigation options
Code
Issues
Pull requests
Actions
Projects
Security and quality
Insights
Commits
Branch selector
master
User selector
All users
All time
Commit history
Commits on Jul 19, 2026
Merge remote-tracking branch 'upstream/master'
gary149
committed
f473a1a
View commit details
Copy full SHA for f473a1a
Browse repository at this point
Commits on Jul 18, 2026
model: rotate injected K/V cache for DFlash (#25823)
Show description for 571d0d5
ruixiang63
and
ggerganov
authored
571d0d5
View commit details
Copy full SHA for 571d0d5
Browse repository at this point
llama-quant : exclude i32 ffn_gate_tid2eid routing table from quantization (#25787)
Show description for 4937ca8
devYRPauli
authored
4937ca8
View commit details
Copy full SHA for 4937ca8
Browse repository at this point
Commits on Jul 17, 2026
opencl: load and use `kernel_gemm_moe_q6_k_f32_ns` from bin kernel lib (#25797)
lhez
authored
86a9c79
View commit details
Copy full SHA for 86a9c79
Browse repository at this point
opencl: read/write MoE dp4a activation tiles to local memory as 128-bit (vectorized LD/ST perf opt) for Adreno GPUs (#25810)
Show description for 6bdd77f
wanghqc
authored
6bdd77f
View commit details
Copy full SHA for 6bdd77f
Browse repository at this point
opencl: transpose q4_K noshuffle scales for coalesced reads (#25805)
wanghqc
authored
86d86ed
View commit details
Copy full SHA for 86d86ed
Browse repository at this point
sync : ggml
ggerganov
committed
7d56da7
View commit details
Copy full SHA for 7d56da7
Browse repository at this point
ggml : bump version to 0.17.0 (ggml/1568)
ggerganov
committed
3727404
View commit details
Copy full SHA for 3727404
Browse repository at this point
tests : initialize all tensors in test_dsv4_hc to avoid NaNs in sentinel tensors (#25822)
Show description for 5d5306b
fairydreaming
and
sszymczy
authored
5d5306b
View commit details
Copy full SHA for 5d5306b
Browse repository at this point
common : auto-download dflash- and eagle3- HF sidecars (#25811)
Show description for 635cdd5
ggerganov
authored
635cdd5
View commit details
Copy full SHA for 635cdd5
Browse repository at this point
ggml-blas: default hadamard mul_mat to cpu routine (#25710)
Show description for 11fd0a6
taronaeo
authored
11fd0a6
View commit details
Copy full SHA for 11fd0a6
Browse repository at this point
vulkan: Support Q2_0 (#25430)
Show description for 788e07d
jeffbolznv
authored
788e07d
View commit details
Copy full SHA for 788e07d
Browse repository at this point
sycl: fix row calculation when K_QUANTS_PER_ITERATION is 1 (#25690)
Show description for 0bd0ec6
malsbat
authored
0bd0ec6
View commit details
Copy full SHA for 0bd0ec6
Browse repository at this point
opencl: add ABS op (#25115)
Gezahegne
authored
b85833e
View commit details
Copy full SHA for b85833e
Browse repository at this point
Commits on Jul 16, 2026
opencl: loads quants as uint for q4_K and q5_K flat mv (optimization for Adreno A7x GPUs) (#25780)
Show description for e8f19cc
wanghqc
and
lhez
authored
e8f19cc
View commit details
Copy full SHA for e8f19cc
Browse repository at this point
docs: added a note about using OpenCl with Adreno 810 (#25786)
akleine
authored
ac2557c
View commit details
Copy full SHA for ac2557c
Browse repository at this point
DeepseekV4: Add fused hyper-connection ops (#25585)
Show description for 0dc74e3
am17an
authored
0dc74e3
View commit details
Copy full SHA for 0dc74e3
Browse repository at this point
hexagon: L2 cache handling rework (dirty bit tracking with lazy flushing) and more MUL_MAT updates (#25762)
Show description for b2dd28a
max-krasnyansky
authored
b2dd28a
View commit details
Copy full SHA for b2dd28a
Browse repository at this point
kleidiai: Add SME vs SME2 distinction in kernel dispatch (#25478)
Show description for f15bd60
matcraje
authored
f15bd60
View commit details
Copy full SHA for f15bd60
Browse repository at this point
vulkan: when using transfer queue for async copies, sync on event_wait to avoid race (#25229)
0cc4m
authored
b15ca93
View commit details
Copy full SHA for b15ca93
Browse repository at this point
conversion: accept BitNetForCausalLM architecture name (#25769)
Show description for 3278e92
khashayarghafouri
authored
3278e92
View commit details
Copy full SHA for 3278e92
Browse repository at this point
TP: fix Phi3, Bert, Plamo2/3, ChatGLM (#25536)
JohannesGaessler
authored
2e1fd76
View commit details
Copy full SHA for 2e1fd76
Browse repository at this point
vendor: update BoringSSL to 0.20260713.0 (#25624)
cabelo
authored
86b719b
View commit details
Copy full SHA for 86b719b
Browse repository at this point
tests: actually exercise `test-recurrent-state-rollback` (#25758)
am17an
authored
32e789f
View commit details
Copy full SHA for 32e789f
Browse repository at this point
server : allow text-only slot save/restore with mtmd (#25076)
CHIPMUNK-T0T
authored
a8dc0e3
View commit details
Copy full SHA for a8dc0e3
Browse repository at this point
convert : fix dflash target tokenizer mismatch during conversion (#25733)
Show description for a55a8c5
ruixiang63
authored
a55a8c5
View commit details
Copy full SHA for a55a8c5
Browse repository at this point
CUDA: Support CUDA Virtual Devices (#25228)
Show description for 79bba02
anavp-nvidia
authored
79bba02
View commit details
Copy full SHA for 79bba02
Browse repository at this point
Enable CUDA graphs on volta+turing (#25749)
heislera763
authored
3f08ef2
View commit details
Copy full SHA for 3f08ef2
Browse repository at this point
server: Ignore empty / non-existing `Origin` headers (#25756)
Show description for 8ee54c8
sdroege
authored
8ee54c8
View commit details
Copy full SHA for 8ee54c8
Browse repository at this point
ggml-cuda : restore prop.integrated on HIP builds (#24233)
Show description for c7d8722
liminfei-amd
authored
c7d8722
View commit details
Copy full SHA for c7d8722
Browse repository at this point
CUDA: dedup MoE gate/up activation quantization (#25441)
Show description for 5839ba3
praneshgo
authored
5839ba3
View commit details
Copy full SHA for 5839ba3
Browse repository at this point
ci : add official website link to release notes (#25728)
Show description for a320cbf
ggerganov
authored
a320cbf
View commit details
Copy full SHA for a320cbf
Browse repository at this point
quant : allow using manual tensor types with --pure (#25716)
ggerganov
authored
56d6e9d
View commit details
Copy full SHA for 56d6e9d
Browse repository at this point
opencl: disable FA and MoE weights repack to work around compiler issues for Adreno 850 GPU (#25745)
Show description for 3dafb58
wanghqc
and
lhez
authored
3dafb58
View commit details
Copy full SHA for 3dafb58
Browse repository at this point
cuda: extract Q1_0 elements via __byte_perm (#25628)
dfriehs
authored
602f828
View commit details
Copy full SHA for 602f828
Browse repository at this point
Previous
Next
You can’t perform that action at this time.