Sign in to view Tuhin Kanti’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Tuhin Kanti’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
West Bengal, India
Sign in to view Tuhin Kanti’s full profile
Tuhin Kanti can introduce you to 1 people at visadb.io
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
297 followers
294 connections
Sign in to view Tuhin Kanti’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Tuhin Kanti
Tuhin Kanti can introduce you to 1 people at visadb.io
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Tuhin Kanti
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Tuhin Kanti’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Websites
- Portfolio
-
https://github.com/tuhinpal
- Personal Website
-
https://thetuhin.com
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
297 followers
-
Tuhin Kanti Pal shared this💪🚀Tuhin Kanti Pal shared thisI have 3 VCs in my inbox :) for Bookeeping.ai, AI Accountant Paula. No disrespect but already bootstraped and profitable :)
-
Tuhin Kanti Pal posted thisAccess to ad free internet should not be expensive. So I built a lightweight DNS ad blocker on ESP32, inspired by Pi-hole, for low cost setups and resource constrained devices. Open source here 👇 #ESP32 #IoT #OpenSource #AdBlocker #DNS #PiHole #Networking #EmbeddedSystems #LowPower #TechForGood #EdgeComputing #Microcontrollers #DIYTech #InternetPrivacy #CyberSecurity
-
Tuhin Kanti Pal reposted thisTuhin Kanti Pal reposted thisI am looking to hire someone who can work with non-profits. We have built an AI bookkeeping product for non-profits :) anyone you know?
-
Tuhin Kanti Pal posted this🚀 Just released: Universal Asymmetric Cryptography Utils A TypeScript library for RSA encryption that works everywhere - Node.js, browsers, Cloudflare Workers, Deno, and more! 🔐 Features: • One API, all platforms • RSA key generation & encryption/decryption • CryptoKey + PEM support • Full TypeScript types ✅ Perfect for: • Exchange symmetric keys • Secure message exchange • Cross-platform data protection • End-to-end encryption Built with Web Crypto API for maximum security and compatibility. npm install asymmetric-cryptography-data-exchange-utils #OpenSource #TypeScript #WebCrypto #JavaScript #Security
-
Tuhin Kanti Pal shared thisDelightful in-browser native PDF editor! Made to fill tax forms at bookeeping.ai. We also added custom implementation of image and text adding, and form filling in PDFs. w Danish Soomro,
-
Tuhin Kanti Pal shared thisMe and my whole team switched from both ChatGPT & Claude to Jan.ai . We are still using Claude and OpenAI's api. Huge cost saving ($40 per user) with more productivity and privacy.w Danish Soomro,
-
Tuhin Kanti Pal shared this5,101 GitHub contributions in 2024 ~1M lines of code!
-
Tuhin Kanti Pal shared thisTwo years ago at visadb.io we built a visa widget. Today, I randomly visited that page after almost 1.5 years and found: - Everything is working - All data is in sync with real data, regulations, and latest events - 500+ websites are using the widget - Latency is under 100ms on every page This feels awesome! w Danish Soomro,
-
Tuhin Kanti Pal shared thisFeeling thrilled & proud to be a part of its active development. Let's solve bookkeeping! 😎Tuhin Kanti Pal shared thisLast year I could not get a tax professional, the one I got asked me $2000, IRS deadline was approaching. So I went on Fiverr and found someone for a few hundred, however, the person made so many mistakes and was not engaged at all. At that time I decided to build AI Bookkeeping assistant for Entrepreneurs and Small businesses. It completes financial tasks via chat like prepping ledgers, statements, invoices and all. Last 6 months my team & I have been working day and night. Finally, we are launching a private beta. Go apply :) https://bookeeping.ai
-
Tuhin Kanti Pal reacted on thisTuhin Kanti Pal reacted on thisI have 3 VCs in my inbox :) for Bookeeping.ai, AI Accountant Paula. No disrespect but already bootstraped and profitable :)
-
Tuhin Kanti Pal reacted on thisTuhin Kanti Pal reacted on thishttps://lnkd.in/grcvsjfW Tired of setting up the same project from scratch every single time? 🥱 Pick your stack. Download your starter. Open your editor. That's it. No more copy-pasting configs from old projects. No more "wait, how did I set this up last time?" No more wasting 2 hours before writing a single line of actual code. Your tech stack. Your custom templates. Your rules — ready in seconds. And if you want to start from scratch? Bring your own code file. Zero repeats. Zero friction. The boring part should be automated. So you can focus on the part that actually matters — building. 🚀 💬 How much time do you waste on project setup every month? Drop it below 👇 #buildinpublic #devtools #webdev #developerexperience #productivity #javascript #opensource #developers #coding #softwareengineerin
-
Tuhin Kanti Pal reacted on thisTuhin Kanti Pal reacted on thisI am looking to hire someone who can work with non-profits. We have built an AI bookkeeping product for non-profits :) anyone you know?
Experience & Education
-
visadb.io
********* ****
-
**** * ** ****** ***** *******
********* ****
-
***** ************
********** ****** ******** ******* 8.0 (CGPA)
-
-
* * * ********* ** *********** * **********
******* ********** *********** *** *****
-
View Tuhin Kanti’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Licenses & Certifications
Projects
-
Serverlesstalent
-
See projectA serverless developer marketplace. This was my second project at visadb.
-
Dynamic Image
-
See projectDynamically generate images for open-graph or website. Using to provide Open Graph images in Visadb at a scale.
Languages
-
English
Full professional proficiency
-
Hindi
Full professional proficiency
-
Bengali
Native or bilingual proficiency
View Tuhin Kanti’s full profile
-
See who you know in common
-
Get introduced
-
Contact Tuhin Kanti directly
Other similar profiles
-
Hardik P.
Hardik P.
Shastri Swami Shree Dharmajivandasji Institute Of Information Technology
29K followersGreater Ahmedabad Area
Explore more posts
-
Anurag(Anu) Karuparti
Microsoft • 36K followers
𝟕 𝐋𝐋𝐌 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐏𝐚𝐫𝐚𝐦𝐞𝐭𝐞𝐫𝐬 Understanding LLM generation parameters is critical for controlling output quality. Here are common generation parameters you may see across LLM APIs: 1. MAX_TOKENS Limits the number of tokens the model generates. Use when: Need to control response length Example: Customer support reply should be no more than 100 words. 2. TEMPERATURE Controls creativity; lower values = focused, higher values = more creative. Use when: • 0-0.3: Factual, more deterministic outputs • 0.7-1.0: Balanced creativity • 1.0-2.0: Maximum creativity Example: Set low for compliance answers, higher for ad copy. 3. TOP_P Sets the probability threshold for token diversity. Higher = more variable. Use when: Need to control diversity while maintaining quality Example: You want five versions of a product description that stay mostly on brand but do not all sound the same. 4. TOP_K Limits the number of tokens considered when predicting the next token. Use when: Want to restrict vocabulary size Example: You are generating simple NPC dialogue for a game and want the language to stay tight, predictable, and easy to control. 5. FREQUENCY_PENALTY Reduces repeated tokens, encouraging more diversity in the response. Use when: Want to avoid repetitive language Example: You are generating a long article or LinkedIn post and want to avoid the model repeating the same phrases again and again. 6. PRESENCE_PENALTY Discourages reusing tokens and encourages generating new ones. Use when: Want novel vocabulary and ideas Example: You are brainstorming campaign concepts and want the model to keep exploring fresh angles instead of staying on the same theme. 7. STOP Specifies when the model should stop generating content. Use when: Need precise stopping conditions Example: You are generating a formatted template or structured output and need the response to end exactly before the next section marker, such as "END" or "###" WHEN TO USE EACH PARAMETER 1. Length Control: • max_tokens: Hard limit on response length 2. Creativity Control: • temperature: Overall randomness (0 = deterministic, 2 = creative) • top_p: Probability threshold for diversity • top_k: Token count limit 3. Diversity Control: • frequency_penalty: Reduce repetition of tokens • presence_penalty: Encourage new tokens 4. Stopping Control: • stop: Define stopping conditions THE COMBINATION STRATEGIES Factual Outputs: • temperature: 0-0.3 • top_p: 0.1-0.3 • frequency_penalty: 0 • presence_penalty: 0 Creative Content: • temperature: 0.7-1.2 • top_p: 0.8-0.95 • frequency_penalty: 0.5-1.0 • presence_penalty: 0.5-1.0 Code Generation: • temperature: 0-0.2 • top_p: 0.1 • frequency_penalty: 0 • max_tokens: as needed Conversational: • temperature: 0.5-0.8 • top_p: 0.9 • presence_penalty: 0.6 • frequency_penalty: 0.3 Master these parameters to control output precisely. Which parameter combination does your use case need?
122
45 Comments -
AJAL RC
Wellmark Blue Cross and Blue… • 553 followers
𝗔𝗜 𝗮𝗴𝗲𝗻𝘁𝘀 𝗰𝗮𝗻 𝘄𝗿𝗶𝘁𝗲 𝗰𝗼𝗱𝗲, 𝗶𝗻𝘃𝗲𝘀𝘁𝗶𝗴𝗮𝘁𝗲 𝗶𝗻𝗰𝗶𝗱𝗲𝗻𝘁𝘀, 𝗮𝗻𝗱 𝗮𝘂𝘁𝗼𝗺𝗮𝘁𝗲 𝘄𝗼𝗿𝗸𝗳𝗹𝗼𝘄𝘀. 𝗕𝘂𝘁 𝘁𝗵𝗲𝘆 𝗵𝗮𝘃𝗲 𝗮 𝗺𝗲𝗺𝗼𝗿𝘆 𝗽𝗿𝗼𝗯𝗹𝗲𝗺. Most LLM-powered systems are stateless. Everything an agent “remembers” during a task is stored in the context window. Think of it like RAM in a computer — temporary memory. Once the session ends or the context window fills up, that memory disappears. That’s fine for chatbots. But what happens when an agent generates something real? A Python script, a remediation playbook, an analysis report, or even a workflow automation. Where does that work actually go? Solutions like Retrieval Augmented Generation (RAG) help agents read external knowledge from vector databases or documents. But RAG is fundamentally read-focused. It helps bring information into the model. It doesn’t solve the right problem. 𝗧𝗵𝗮𝘁’𝘀 𝘄𝗵𝗲𝗿𝗲 𝗮𝗴𝗲𝗻𝘁𝗶𝗰 𝘀𝘁𝗼𝗿𝗮𝗴𝗲 𝗰𝗼𝗺𝗲𝘀 𝗶𝗻. Agentic storage gives agents persistent memory for their outputs — almost like giving an AI agent a hard drive rather than just RAM. Many architectures are increasingly relying on the Model Context Protocol (MCP), which provides a standardized interface for agents to interact with tools, storage systems, and APIs. Instead of building custom integrations for every system, agents can call standardized tools exposed through an MCP server. But giving autonomous agents write access introduces risk. 𝗦𝗼 𝘀𝗲𝗰𝘂𝗿𝗶𝘁𝘆 𝗹𝗮𝘆𝗲𝗿𝘀 𝗮𝗿𝗲 𝗲𝘀𝘀𝗲𝗻𝘁𝗶𝗮𝗹: • Immutable versioning so every change can be rolled back • Sandboxing so agents operate only in controlled environments • Intent validation to verify why an agent is performing high-impact actions As AI moves from chat interfaces → autonomous agents, infrastructure will need to evolve too. 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝘀𝘁𝗼𝗿𝗮𝗴𝗲 𝗺𝗶𝗴𝗵𝘁 𝗯𝗲𝗰𝗼𝗺𝗲 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗸𝗲𝘆 𝗯𝘂𝗶𝗶𝗹𝗱𝗶𝗻𝗴 𝗯𝗹𝗼𝗰𝗸𝘀 𝗼𝗳 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻 𝗔𝗜 𝘀𝘆𝘀𝘁𝗲𝗺𝘀. Curious how others are thinking about persistent memory and storage for AI agents. #AI #AgenticAI #AIInfrastructure #LLM #SoftwareEngineering #AIEngineering #SystemDesign
10
4 Comments -
Raj Kapadia
Insight Global • 2K followers
🚨 An Indian company just claimed to outperform GPT-5.2, Gemini, and Mistral in OCR. Servum recently introduced Servum Vision, and according to their published results, it beats major models on the OLM OCR Benchmark — a dataset designed to test: • Table parsing • Reading order • Mathematical formula rendering • Column alignment • Complex PDF structures They also claim support for 22 Indic languages, which is huge if true. I decided to test it myself. Here’s what I found: ✅ Performs surprisingly well on handwritten notes ✅ Handles cursive English fairly accurately ✅ Strong table structure understanding ⚠️ Benchmark GitHub repo currently missing One interesting point — the API code (Python/TypeScript) seems partially implemented. So is this a breakthrough in OCR? Possibly. But transparency in benchmarking will be key. I’ve shared a full technical breakdown and demo on YouTube. India building strong AI infra is something I’m genuinely excited about. Would love to hear your thoughts — especially from people working in Document AI & OCR systems. #AI #OCR #GenerativeAI #IndianStartups #MachineLearning #DocumentAI
3
3 Comments -
Anshul Kumbhare
KPIT • 5K followers
“Perplexity AI Just Beat ChatGPT, What That Really Means” In a surprise move, Perplexity AI has become the #1 most downloaded app across Android and iOS in India, leapfrogging ChatGPT, Google Gemini, and even WhatsApp. But this isn’t just a chart-topping moment. It’s a signal. 📈 The Rise of AI-Powered Search Perplexity isn’t just another chatbot. It’s a search engine reimagined, combining real-time web results with conversational answers. Users aren’t just asking questions. They’re getting context-rich, citation-backed insights in seconds. 🔍 Why This Matters for Data Scientists, Marketers, and Founders SEO is changing: Traditional keyword stuffing is out. AI-generated, user-centric content is in. Search behavior is shifting: People want answers, not links. Trust is the new currency: Perplexity’s citation-first model builds credibility fast. 🧠 What You Should Do Now Start optimizing for AI-native search platforms. Focus on clarity, citations, and conversational tone in your content. Think beyond Google, your next customer might find you through Perplexity. The future of search isn’t just smarter, it’s more human. Are you ready to be found?
2
-
Tushar Balki
Varaha • 939 followers
𝗦𝗲𝗲𝗶𝗻𝗴 𝗹𝗼𝘁𝘀 𝗼𝗳 "𝗦𝗮𝗿𝘃𝗮𝗺 𝗯𝗲𝗮𝘁𝘀 𝗖𝗵𝗮𝘁𝗚𝗣𝗧/𝗚𝗲𝗺𝗶𝗻𝗶" 𝗵𝗲𝗮𝗱𝗹𝗶𝗻𝗲𝘀 𝗹𝗮𝘁𝗲𝗹𝘆? That's missing the real story. 👇 When Sarvam performs strongly against global models like ChatGPT, Gemini, DeepSeek, or Grok in Indian language tasks, it's not about "beating global AI" overall. It's a clear signal about strategic positioning. 𝗦𝗮𝗿𝘃𝗮𝗺 𝘄𝗶𝗻𝘀 𝘄𝗵𝗲𝗿𝗲 𝗴𝗹𝗼𝗯𝗮𝗹 𝗺𝗼𝗱𝗲𝗹𝘀 𝗱𝗼𝗻'𝘁 𝗽𝗿𝗶𝗼𝗿𝗶𝘁𝗶𝘇𝗲. Frontier models chase broad global generality (English-heavy, clean inputs, universal tasks). An India-focused model like Sarvam optimizes for the realities that matter here: Deep multilingual nuance (beyond basic translation) Code-mixed conversations (Hinglish, Tanglish how Indians actually speak online) Voice-first users and regional accents Public sector workflows (govt forms, scanned docs, tables/math) Data localization and regulatory alignment That's not a smaller ambition. It's a different and crucial layer of the stack. The common mistake: thinking global = superior and local = limited. Remember: • Amazon started with US e-commerce. • Tencent started with China chat. • Grab started in Southeast Asia mobility. • Local depth can turn into real infrastructure power. If Sarvam is only chasing benchmark optics, it won't last. But if it's building linguistic infrastructure for 1.4B people across 20+ languages that's not a niche. That's a moat. The bigger shift we're seeing: 𝗔𝗜 𝗶𝘀𝗻'𝘁 𝗷𝘂𝘀𝘁 𝗮 𝗿𝗮𝘄 𝗺𝗼𝗱𝗲𝗹 𝗿𝗮𝗰𝗲 𝗮𝗻𝘆𝗺𝗼𝗿𝗲. 𝗜t'𝘀 𝗮 𝗱𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗶𝗼𝗻 + 𝗰𝗼𝗺𝗽𝗹𝗶𝗮𝗻𝗰𝗲 + 𝗹𝗼𝗰𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝗿𝗮𝗰𝗲. Winners may not be the loudest but the most embedded #AIinIndia #SovereignAI #MultilingualAI #AIIndia https://lnkd.in/dRy2SxKK
6
2 Comments -
Prashanth T
primesoma • 516 followers
I tested seven AI models, ranging from 0.8B to frontier, on the same challenging problem. This was not a benchmark; it involved a real multilingual parsing task from production, requiring code-switched numbers in Tamil, Hindi, and English at the same time. The results shattered all my assumptions. The 0.8B model stripped away all non-numeric characters, returned None for everything, and blamed Python for not supporting Tamil. It was clearly broken. The 4B model created something limited. It understood exactly what was wrong and documented every failure honestly, with no hallucinations. The 9B model mistakenly added a Tamil verb meaning "write" to a number dictionary as the word for seven. It had incorrect values everywhere and experienced silent failures. The 30B model (GLM-4.7) invented an entire Tamil transliteration system that doesn't exist, duplicated dictionary keys, and failed to parse the original test variables. The 397B model MoE (Qwen) made up a Tamil word that isn’t real and hardcoded it, labeling it as a "Code-Switch Hybrid Specific" in its documentation. The GLM-5 model, priced at $1.55 per million parameters, caught a bug that the $10 per million model missed. It interpreted "Nineteen eighty-four" as 103 instead of 1984, which was both silent and catastrophic. The Claude Opus model, also at $10 per million, demonstrated the deepest understanding of language. It recognized Tamil sandhi morphology and traced Hindi's roots back to Prakrit but missed the practical bug. Here’s what the data shows: - The most honest model had 4B parameters. - The most dangerous failure came from the 397B model. - The cheapest frontier model caught what the most expensive one missed. Confidence doesn’t scale with size; it simply learns to hide ignorance better. The only way to find out where your model fails is to test it on your actual problem. Have you done that?
17
-
Mukesh Karn
Emorphis Technologies • 8K followers
🚀 Google Gemini 3 — Launched Nov 18, 2025: A Big Leap in Reasoning & Agentic AI 🚀 Google just shipped Gemini 3, and it’s being billed as a major step forward — not just bigger, but smarter. Here’s the short, sharp breakdown you need (1-min read): 🔎 Core Capabilities Advanced Reasoning (Deep Think) — designed to reason through multi-step scientific, analytical and mathematical problems (Deep Think mode rolling out to Ultra users). Native Multimodality — text, images, audio, video and code in one model; can analyze long videos or whole codebases in a single prompt. Huge Context Window — ~1 million tokens, so it can hold entire books, long repos, or hours of video in context. 💻 Coding & Agentic Workflows Vibe Coding — understands the intent and “vibe” of a project (less rigid prompting for building complex apps). Google Antigravity — an agentic dev environment where Gemini 3 can act autonomously: write code, run terminal tests, and inspect results in a browser workflow. ⚡ Performance & Benchmarks (per Google) Claims to beat leading rivals on major reasoning benchmarks (e.g., top ranks on LMArena). Reported high scores on specialized exams — showing stronger multi-step reasoning and domain expertise. 🔗 Integration & Access Generative UI in Search — produces interactive mini-apps (graphs, planners) directly in search results. Deeper Workspace integration — Docs, Drive, Gmail workflows (e.g., “Read 50 PDFs and summarize into a spreadsheet”). India note: Reliance Jio announced a partnership offering free Gemini 3 Pro access for 18 months to eligible Jio Unlimited 5G users. 🧩 Model Tiers (what to expect) Gemini 3 Pro — high-capability model available now. Gemini 3 Deep Think — specialized reasoning variant (rolling out soon). Gemini 3 Flash — likely a faster, lighter edition for high-throughput tasks. 🤔 Why this matters: Gemini 3 pushes beyond “better text” — Google is betting on reasoning, multimodal understanding, and agentic workflows. If the claims hold up in real-world use, this will reshape how we build, debug, and ship products with AI assistance. What excites you most — the Deep Think reasoning, agentic coding, or the massive context window? 👇 ↫↫↫↫↫ 🌐 Click the "Follow" button now and let's embark on this incredible professional journey together. Don’t miss out on the latest updates, valuable resources, and exciting discussions. ↬↬↬↬↬ 👉🏻 𝗙𝗼𝗹𝗹𝗼𝘄 Mukesh Karn 👨💻 𝗳𝗼𝗿 𝗺𝗼𝗿𝗲. 🧡 𝗛𝗶𝘁 𝗹𝗶𝗸𝗲, 𝗶𝗳 𝘆𝗼𝘂 𝗳𝗼𝘂𝗻𝗱 𝗶𝘁 𝗵𝗲𝗹𝗽𝗳𝘂𝗹!! 🔖 𝗦𝗮𝘃𝗲 𝗶𝘁 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗳𝘂𝘁𝘂𝗿𝗲 📤 𝗦𝗵𝗮𝗿𝗲 𝗶𝘁 𝘄𝗶𝘁𝗵 𝘆𝗼𝘂𝗿 𝗰𝗼𝗻𝗻𝗲𝗰𝘁𝗶𝗼𝗻𝘀 💭 𝗖𝗼𝗺𝗺𝗲𝗻𝘁 𝗿𝗶𝗴𝗵𝘁 𝗻𝗼𝘄 — which Gemini 3 feature would change your workflow? #GoogleGemini3 #AI #GenerativeAI #MachineLearning #AgenticAI #Multimodal #DeepThink #VibeCoding #AIforDevelopers #TechNews #Gemini3
14
2 Comments -
Alok Nayak
The Interview World • 20K followers
“Which API do I call?” That question used to signal technical competence. In the LLM era, it signals outdated thinking. Large Language Models are flipping the interface stack on its head. The real work is no longer choosing endpoints or memorizing function signatures, it’s understanding intent. What does the user actually want to get done? Natural language is becoming the new control plane. Not as a novelty, but as infrastructure. Protocols like Model Context Protocol (MCP) allow systems to translate human goals into coordinated action across tools, data, and workflows, without forcing humans to think like machines. This shift is profound for enterprises. Integration becomes orchestration. APIs become capabilities. And software design moves from “How do I call this?” to “What outcome should this enable?” The winners won’t build better APIs. They’ll build clearer intent models, smarter orchestration, and systems that meet humans where they are. That’s not a tooling upgrade. That’s a mindset reset. #AI #LLM #EnterpriseAI #SoftwareArchitecture #DigitalTransformation #ProductStrategy #FutureOfWork VentureBeat Dhyey Mavani https://lnkd.in/g2PwWuub
-
Hritik Rai
Founda • 2K followers
Getting grounded answers from small LLMs in RAG without more hardware I had the usual RAG setup: Notion docs, vector search, then an LLM. Retrieval was fine. The problem was what happened after retrieval. The problem A 14B model was getting 5 full document chunks as context (15–20KB of text). It would drift into generic answers instead of sticking to the docs, output JSON instead of natural language, or time out. It wasn't a retrieval problem. The right documents were being selected. The model was drowning in context and couldn't use them. The idea that helped The CLaRa paper (Apple) suggests compressing documents before the answer model instead of passing raw chunks. Their setup uses a trained compression model and latent representations. I couldn't do that. What I had 16GB VRAM, one 14B model already running for chat, and no room for another model or training. I needed something simpler. The solution I used the same 14B model for one extra step at query time. Before sending retrieved chunks to the answer call, I run a few more calls: for each chunk, I ask the model to list only the facts relevant to the user's question as bullets. I collect those bullets into a short "document facts" block and send that to the final answer call instead of the full chunks. About 2–3KB instead of 15–20KB. Same model, a few extra calls per query, no new hardware. The trade-off Latency. Each query now includes several extraction calls (around 10–20 seconds extra). In return, answers stay grounded in the docs instead of hallucinating or breaking into JSON. What to call it I call it "query time extractive compression." "CLaRa inspired" is fair for the idea (compress before answer), but my approach is different: no latent model, no training, just prompting the same LLM to extract facts per doc. Takeaway You don't need the full research implementation to use its core idea. Given your constraints, a simple query time step with your existing model can turn unusable RAG into something that works. If you've dealt with similar RAG + small model problems, I'd like to hear how you solved them. #RAG #LLM #MachineLearning #AI #RetrievalAugmentedGeneration #NLP
31
3 Comments -
Akbar Ali
Headlyne_app • 3K followers
JSON was built for the web. TOON was built for tokens. It’s December 2025. If your LLM applications are still relying solely on raw JSON for heavy structured data lifting, you are burning money on redundant syntax. Repeated keys, endless quotes—it’s bloat that LLMs shouldn't have to process. We needed a format that treats tokens as a scarce resource. We got TOON (Token-Oriented Object Notation). TOON is the "Steve Jobs" approach to data structuring: remove everything non-essential until only the pure signal remains. It’s simpler, more reliable, and just feels right. As the infographic below details, cutting-edge teams have already made the switch for internal agent comms and RAG. The impact is undeniable: 🚀 30–60% fewer input tokens (huge savings on large contexts). 🎯 3–10x fewer parse retries (better structured reliability). 💸 End-to-end cost reduced by 25–45%. JSON will always be the universal default for public APIs. But for high-performance AI workflows, TOON is the new standard. The switch typically pays for itself in weeks. Have you tried TOON yet? Let me know in the comments below. #TOONformat #LLM #GenerativeAI #DataEngineering #SoftwareArchitecture #AI2025
11
5 Comments -
Kurt Cagle
The Cagle Report • 28K followers
This is the next stage of AI. I've been coming to the same conclusion, albeit from a slightly different direction. Have you ever had a long conversation, and in the middle of things something's brought up that you think is really insightful, but then you lose the thread and you can only dimly remember what it was that you talked about? That's the current state of Transformers. Ideas are amorphous, often are highly contextual, and they shift and evolve over time. Contexts are both finite and bounded. Internally we abstract, but abstraction is almost invariably lossy. The next possible stage is persistent memory. RAG by itself isn't enough - it's primarily just a hack of LangChain to allow external services. What we need is a way of growing and holding named in-memory graphs so that they retain persistence. We need addressibility into that graph. Maybe Deepseek's architecture will be the one to do so, maybe someone else will figure out an alternative, but I think that has to happen for language models to get (mostly) past decaying coherence.
46
8 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content