mintyleaf 0d6c3a7d57
feat: include tokens usage for streamed output (#4282)
Use pb.Reply instead of []byte with Reply.GetMessage() in llama grpc to get the proper usage data in reply streaming mode at the last [DONE] frame

Co-authored-by: Ettore Di Giacinto <mudler@users.noreply.github.com>
2024-11-28 14:47:56 +01:00
..
2024-06-23 08:24:36 +00:00
2024-06-23 08:24:36 +00:00
2024-10-30 10:57:21 +01:00
2024-06-23 08:24:36 +00:00
2024-06-23 08:24:36 +00:00