<?xml version="1.0" encoding="UTF-8" ?><!-- generator=Zoho Sites --><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><atom:link href="https://www.aheadcrm.co.nz/blogs/tag/LLM/feed" rel="self" type="application/rss+xml"/><title>aheadCRM - Blog #LLM</title><description>aheadCRM - Blog #LLM</description><link>https://www.aheadcrm.co.nz/blogs/tag/LLM</link><lastBuildDate>Tue, 22 Sep 2026 12:03:48 -0700</lastBuildDate><generator>http://zoho.com/sites/</generator><item><title><![CDATA[Zoho goes all in with AI - bold or inevitable?]]></title><link>https://www.aheadcrm.co.nz/blogs/post/zoho-goes-all-in-with-ai-bold-or-inevitable</link><description><![CDATA[The news On July 17, 2025 Zoho launched Zia LLM and deepened its AI portfolio with agents, an agent builder, MCP support and an agent marketplace. Key a ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_-yKlOOC-QuqljzF5yFe8vQ" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_Vlw2AkMQTu-2X69FlBd2Qw" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_W9wpsRPmSj2mp_my4cBW9w" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_uceKWyC9Q7-NDUAtPweeoQ" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><h1 class="wp-block-heading">The news</h1><p>On July 17, 2025 Zoho <a href="https://www.zoho.com/news/zoho-launches-zia-llm-and-deepens-ai-portfolio-with-prebuilt-agents-custom-agent-builder-mcp-and-marketplace.html?zwcDate=july%2017%2C%202025">launched Zia LLM and deepened its AI portfolio</a> with agents, an agent builder, MCP support and an agent marketplace.</p><p>Key announcements from the press release include:</p><ul class="wp-block-list"><li><strong>In-House LLM</strong>: Zoho has developed its own large language model, Zia LLM, which comes in three sizes (1.3B, 2.6B, and 7B parameters) to optimize for different business use cases. This allows customers to leverage AI while keeping their data within Zoho's ecosystem, ensuring privacy. The three models allow Zoho to always optimize the right model for the right user context, striking the proper balance between power and resource management. This focus on right-sizing the model is an ongoing development strategy for Zoho.</li><li><strong>Speech-to-Text Models</strong>: The company also unveiled two proprietary Automatic Speech Recognition (ASR) models for English and Hindi, with plans to support more languages in the future.</li><li><strong>Prebuilt AI Agents</strong>: To facilitate immediate adoption, Zoho has introduced a range of AI agents that are integrated directly into its products. These agents are designed to automate tasks for various business roles such as sales development, customer support, and account management.</li><li><strong>Global and Private Cloud Deployment</strong>: The new Zia LLM will be deployed across Zoho's data centers in the US, India, and Europe.</li><li><strong>Continued Support for Other Models</strong>: While promoting its own AI, Zoho will continue to support integrations with other popular large language models like ChatGPT, Llama, and DeepSeek.</li><li>Zoho will continue to scale Zia LLM’s mode sizes. A2A capabilities are on the roadmap.</li></ul><h1 class="wp-block-heading">The bigger picture</h1><p>Enterprise software has been a platform game for a long time. AI, in particular generative and agentic AI, have upped the ante in this respect. The competition has become even more one that is based around platforms. An enterprise software vendor that wants to address a significant portion of an enterprise’s value chain cannot go without agents, an agent builder, and a marketplace anymore. This is true whether they use an own AI, a derivative of an open source one, or build it using their own stack.</p><p>Plus, support of MCP becomes non-negotiable as most customers run on more than one platform. A2A support then is the next step. Both can also be used as a bridge head and Trojan horse to work towards an increased share of the customers’ value chains.</p><p>At the same time this AI battleground is evolving with an enormous pace which seems to accelerate every day. This makes it important for buyers to have a deep look at who to engage with while vendors need to offer a strong narrative around a compelling value proposition.</p><h1 class="wp-block-heading">My point of view and analysis</h1><p>With this announcement it is clear that Zoho is serious about AI. With this, the company raises its AI profile significantly, particularly in the CX space, as Zoho CRM is its most mature product.</p><p>The company works on AI capabilities, tooling and infrastructure tor quite some time now. Plus, it owns its hardware and data centers. This includes already having deployed more than 80 AI algorithms in the past.</p><p>Now it also has a family of LLMs – a family that is poised to grow, initially with a 32bn parameter model. And frankly, to me this is the biggest splash although all other parts of the announcement show a round offering.</p><p>One of the core questions that I asked myself is whether the world needs another LLM or whether it would be sufficient to adapt an existing open source one to own needs. On the other hand, this would not be the Zoho way. Zoho is adamant in staying in control of its own destiny. As part of this, the company is convinced that it needs to own its software stack, as only this way gives full control, full understanding of its behavior and all options. This is certainly also true for an LLM. As Raju Vegesna, Chief Evangelist of Zoho, said: “<em>If AI is going to be the future, and if you are a technology company, what’s more important than creating that core part of technology that is going to define its future?</em>”</p><p>In addition, Zoho’s LLM family comes with a twist. It is not addressing consumer tasks but is specialized on what Zoho customers require. And Zoho customers are predominantly B2B businesses. In that sense, and true to the objective of delivering “contextual AI”, Zoho follows a route that is similar to the one that SAP takes with its own foundation model. In contrast to SAP, Zoho did not train the model on customer data (SAP trains its foundation model with customer consents) which is in line with the company’s strong stance on privacy and trust. Offering differently sized models is all about “rightsizing”, i.e., the economic and performant use of valuable compute resources. In essence, Zoho offers to not bring a sledgehammer to a job when a mallet is sufficient. If the system is able to select the right model automatically, this can be a game changer.</p><p>The second really interesting part of the announcement is about its automatic speech recognition (ASR) models. This is again a matter of being independent. Further, a well-functioning ASR can help businesses raise significant efficiencies, if the models work well. As per Zoho’s tests they compare well against larger models. The current limitation is that the ASR models support only English and Hindi at the moment. I expect further widely used languages being supported soon.</p><p>The offering of prebuilt agents, an MCP server, an agent studio and a marketplace show that Zoho is serious, but they are also table stakes these days. They are the foundation for scale. These offerings, in combination with the upcoming A2A support, open up the ecosystem game by enabling partners to innovate fast and to facilitate getting into new accounts by being open.</p><p>All in all, this is a bold, yet not unexpected move by Zoho. It helps the company to strengthen its enterprise credibility, while maintaining the needs of SMBs squarely in the sights. Especially considering Zoho’s strong stance on value and pricing, the company now places significant pressure on the competition but also on itself, as it is still testing whether it can deliver these capabilities as part of the subscription or whether at least parts of it need to get priced. But just doing this exercise with the goal of delivering these capabilities as part of the package. Even if Zoho will not be able to absorb the cost of running models and agents, pricing can be expected to be extremely competitive.</p><p>As Raju Vegesna said, it is a marathon.</p><p>And this marathon has only just started.</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Thu, 24 Jul 2025 11:17:58 -0400</pubDate></item><item><title><![CDATA[LLM Showdown: Comparing ChatGPT, Gemini, and Grok for Automated News Research]]></title><link>https://www.aheadcrm.co.nz/blogs/post/llm-showdown-comparing-chatgpt-gemini-and-grok-for-automated-news-research</link><description><![CDATA[The analyst’s day is full of research. Now, this is the age of AI and AI is here to help, isn’t it? As everyone is talking about copilots and AI agent ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_aA1-EemuS-CA0kel7TE6Ew" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_-C0eWl7DT7GK9AIeiEYPLg" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_zlR6LRDSSi2eoVWDICjKpg" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_is9GdXmaT3SG6hl0h32HIQ" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><p>The analyst’s day is full of research. Now, this is the age of AI and AI is here to help, isn’t it? As everyone is talking about copilots and AI agents, why not using the tools at hand to do a little research on research.</p><p>NB., no one really has a good definition of an AI agent, so this might become an additional topic for research.</p><p>But I digress.</p><p>Imagine the following project at hand, which is not only interesting for analysts, btw, but also for a variety of roles in the corporate world. Let’s call it vendor (competitor) monitoring. The job is the following:</p><ul class="wp-block-list"><li>Research reputable sites for news about a number of vendors, relating to a set of keywords. Reputable sites are high quality news sites, high quality tech publications, high quality analyst sites and, of course the news pages of the vendors in question.</li><li>Limit the time frame of the search matching to the cadence of my information requirement, e.g., “yesterday” for a daily update or “last week” for a weekly update</li><li>Provide a summary of the news</li><li>Give an assessment of how the news affects the positions of the vendors in the marketplace re the key words in question</li><li>Provide these news with their assessments as a prioritized list, sorted from high impact to low impact</li><li>Add an executive summary as a preface</li><li>Send it to me as an email</li></ul><p>So far, so simple. After all, a lot of folks, yours truly included, do this every day. And it is taking quite some time. So, this job is a perfect one for an automated update beyond a CSS feed. And it seems like a perfect job for an LLM turned agent – or is it a copilot?</p><p>Now, the basic question is: Which one to use? After all, there are plenty, from free to not so free ones. Answering this question turns into yet another interesting experiment: Why not ask some LLMs for their evaluation of suitability? Kinda meta, but an interesting one.</p><p>So, I did just that: I asked ChatGPT 4.5, Grok 3, and Gemini in its 4 versions 2.0 Flash Thinking Experimental, 2.0 Flash, 1.5 Pro with Deep Research, 2.0 Pro Experimental for their analysis of which of them is best suited for the research task at hand.</p><p>For this, I used the following, simple prompt:</p><p>describe the different capabilities and limitations of Gemini 2.0 Flash Thinnking Experimental, 2.0 Flash, 1.5 pro with deep research, 2.0 pro experimental, Grok 3 and chatGPT 4.5, both with and without deep reasoning. Which model is best to support the following use case:&nbsp;&nbsp;&nbsp;</p><p>research the web for news on a given set of companies and a given set of topics. The news shall cover the past 2 days only&nbsp;&nbsp;&nbsp;</p><p>assess the news regarding their impact on the companies' market positions re the given set of topics&nbsp;&nbsp;&nbsp;</p><p>create this in the form of a report&nbsp;&nbsp;&nbsp;</p><p>do this as a daily scheduled task&nbsp;&nbsp;</p><p>accuracy, reasoning and reliability are of high importance. Speed is of lower importance.&nbsp;&nbsp;</p><p>generate a comparison table, give a recommendation and justify the recommendation.</p><p>The results are quite interesting.</p><ul class="wp-block-list"><li>Gemini 1.5 Pro with Deep Research resolves to <strong>Gemini 2.0 Flash</strong> because it offers “<em>a balance of reasoning, accuracy, reliability, and tool use necessary for fulfilling the requirements of the specified use case. Its production-ready status, combined with its ability to handle complex analysis and generate comprehensive reports, makes it the ideal LLM for this task</em><em>.</em>” It looks at Grok as the runner up.</li><li>Gemini 2.0 Pro Experimental recommends <strong>Gemini 1.5 Pro with Deep Research</strong> as it “<em>offers the best balance of accuracy, reasoning, and reliability for your demanding research and reporting task, leveraging Google's strengths in both LLMs and web search.”.</em> This is closely followed by Grok 3.</li><li>Gemini 2.0 Flash Thinking Experimental suggests doing a competitive evaluation of <strong>ChatGPT 4.5</strong> with Deep Reasoning and <strong>Gemini 2.0 Pro Experimental</strong> – as they are a tie. “<em>Which is truly &quot;better&quot; in practice will depend on your specific prompts, data, and tolerance for latency and cost</em>”.</li><li>ChatGPT 4.5 (without Deep Research) suggests going for <strong>ChatGPT 4.5</strong> with Deep Research as the model “<em>offers enhanced reasoning capabilities, reduced hallucination rates, and a broader knowledge base, aligning well with the requirements for accuracy, reasoning, and reliability in daily scheduled tasks.</em><em>&nbsp;</em><em>While models like Gemini 2.0 Flash Thinking Experimental and Grok 3 also provide advanced features, ChatGPT 4.5's maturity and proven track record make it a suitable choice for generating comprehensive and reliable reports.</em>​”.</li><li>Grok 3 in Deep Research mode suggests using <strong>Gemini 2.0 Pro Experimental</strong> as its “<em>advanced reasoning capabilities, as evidenced by its performance in complex tasks and 2 million token context window, make it ideal for researching news, assessing market impact, and generating daily reports (<a href="https://deepmind.google/technologies/gemini/pro/" target="_blank" rel="noreferrer noopener">Gemini 2.0 Pro</a>). The integration with Google Search ensures access to recent news, and as a Google product, it likely offers high reliability for scheduled tasks, aligning with the emphasis on accuracy and reasoning over speed. While Gemini 1.5 Pro with Deep Research is tailored for research, Gemini 2.0 Pro Experimental, being a newer model, likely offers superior capabilities</em>”. Grok looks at Grok as the runner up is it offers advanced reasoning and Deep Search “<em>but potential biases from X data integration”.</em></li></ul><h1 class="wp-block-heading">So, what does this tell me?</h1><p>There is probably a bit of self-serving involved in the LLM’s assessments and suggestions. At least Google consistently suggests a Google model and ChatGPT suggests itself. What is a bit confusing is that</p><p>An interesting side remark is that only Gemini 1.5 Pro with Deep Research, ChatGPT 4.5 and Grok 3 provide the sources used for the research. Perplexity does this, too. Providing references is important for validating the results.</p><p>It looks like the results delivered by these LLMs seem to favor Gemini 2.0 Pro Experimental and ChatGPT 4.5, though, although I am impressed by Grok 3. On the other hand, one needs to know that “experimental” means exactly that – the models are not yet fully stable.</p><p>Having said this, if one needs to perform research tasks, as many of us need to, environment matters. Especially smaller businesses often run Google Workspace. In the case that they subscribed to the Business Standard Edition (like I am doing), Gemini is readily available, there is probably no immediate need to purchase an additional ChatGPT license (I have a pro subscription) or a Grok or Perplexity subscription. This is especially true as most of these tools use a lot of the data that users provide to improve their services, which is especially true for free services. Grok, in its privacy statement explicitly recommends to not input any personal data – as it will be used.</p><p>In summary, if and when I need to do research, I’ll use Google Gemini 1.5 Pro with Deep Research and Gemini 2 Pro Experimental as my preferred option, simply because ChatGPT 4.5 with Deep Research only offers limited runs per month. As it doesn’t cost much, additionally running the same research – potentially with a slightly changed prompt to cater for model differences – I will use ChatGPT and (if no sensitive data involved) Grok 3 in addition. Worst case, this gives me additional food for thought.</p><p>What do you think?</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Wed, 12 Mar 2025 20:22:38 -0400</pubDate></item><item><title><![CDATA[The ABC of Zoho AI]]></title><link>https://www.aheadcrm.co.nz/blogs/post/the-abc-of-zoho-ai</link><description><![CDATA[During ZohoDay24, Zoho amongst other topics, gave some insight into how the company looks at AI. Raju Vegesna presented Zoho’s AI vision and progress. ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_-royAZ0MSTahnzQcQ_0arQ" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_-IE-2ZgaRMuL1XnCx03ppQ" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_sFKwSdarQKuQGd7PhJ5jYg" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_6RlHIiXOTHqTNeMzmIHVzQ" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><p>During ZohoDay24, Zoho amongst other topics, gave some insight into how the company looks at AI. <a href="https://www.linkedin.com/in/rajuvegesna1/">Raju Vegesna</a> presented Zoho’s AI vision and progress. Additionally, I had the opportunity for a one on one with Zoho’s director of AI research, <a href="https://www.linkedin.com/in/ramprakashramamoorthy/">Ramprakash (Ram) Ramamoorthy</a>. If you want to listen and watch the interview, you can do this <a href="https://youtu.be/qdiYEJm-k6w">here</a>.</p><figure class="wp-block-embed is-type-video is-provider-youtube wp-block-embed-youtube wp-embed-aspect-16-9 wp-has-aspect-ratio"><div class="wp-block-embed__wrapper">https://youtu.be/qdiYEJm-k6w</div>
</figure><p>Both represented a vision that is refreshingly differentiated from the current hype with everyone and their dog talking like large language models, LLMs, are the everything one needs.</p><p>Well, let me tell you: They aren’t. But let me come to this point later.</p><p>In addition to not every language model being created equal, and typical for a hype, there is still too much talk about the technology itself, whereas in the words of Raju and Ram the best AI implementation is “<em>when the customer doesn’t know they are using AI but finds value in the output</em>”. This resonates very well with me, as one of my beliefs is that the customer shouldn’t care about the technology that is used to achieve the desired outcome, within some constraints like legality, ethics, and efficiency, of course.</p><p>Zoho is a technology vendor with a focus on business applications. So, Zoho quite quickly realized that consumer type AI that e.g., helps with spell checks, or nowadays research, suffers from two fundamental flaws: lacking privacy/security and accuracy when it comes to business applications. Both violate some of Zoho’s core tenets, namely their pursue of privacy and business applications that offer a lot of value to the customer. Take the example of improving one’s writing – for some years now, this is offered by Zoho Writer, instead of making the user copy and paste potentially sensitive information into an external tool with potentially questionable guardrails.</p><p>Similarly for the use of generative AI, or AI in general, in a business context. LLMs, as they are offered by vendors, regularly miss the business context. This means that their responses to a prompt are less than accurate – they are hallucinating. AI works best in context, in this instance in a business context – your business’s context. This requires a process called grounding, or RAG – retrieval augmented generation.</p><p>Or in plain words: It requires business data and analytics on it.</p><p>In Zohos words, this is contextual intelligence with decision intelligence adding on to it. Decision intelligences creates recommendations and/or actions based upon the findings. Raju Vegesna used a revenue and profit timeline to make this point. Business analytics shows a drop in profit. AI is perfectly able to identify this anomaly. But the real questions are why and what to do. The contextual intelligence correlates this drop to other factors, it gives a diagnosis, and the decision intelligence gives a recommendation on what can (or should?) be done, based on the diagnosis.</p><p>Which brings us back to the topic of large language models not being a silver bullet. Language models, especially multi modal ones, can increase the accuracy of e.g., the OCR of an expense application significantly. Ram and his team observed that, in the case of a receipt only having a merchant logo instead of its name, the accuracy of identifying the merchant “<em>has gone up from somewhere {around] 75% to 98%</em>” because of using a multi-modal models. Even a small language model already increases the OCR accuracy, e.g, by not confusing the cash given with the sales price. Scale this up. This increase in accuracy translates to a far higher efficiency of the expense process. And now look at the end-to-end expense process. It starts by taking a picture of the receipt and then extracts and infers different information. On submission the expense report is analyzed for anomalies and policy violations. This multi-step process can be efficiently handled by a variety of specialized models, from narrow models (OCR), small language models (text extraction and inference), mid-sized models (anomaly detection) to LLMs (policy violation.&nbsp;</p><p>Another example is in the legal field, using electronic contract signature as an example. Again, a multi-step process is followed that leads from phishing detection via a narrow model, possibly translation via a small language model, summarization, and anomaly detection via a medium language model and finally the development of recommendations if the suggested document is not compatible with policies.</p><p>Similarly, a warranty process in customer service.</p><p>All these AI supported processes have two things in common: the user is oblivious to the use of AI and secondly, the orchestrated interaction of smaller models makes the process more resource efficient without endangering process efficiency.&nbsp;</p><p>Ram maintains that “<em>Even though these models are trained using energy intensive CPUs, we are able to run the inference on CPUs, and that that means I'm contributing positively to the enrollment. And again, I'm not passing on the GPU tax to my customer. So, find out models that work for different levels. I mean, it's okay, three to 5 million model is more than enough to identify the name of the merchant on a receipt. And a 50 billion model is enough to rephrase legal statements given its specialized in that domain, that domain and the context is what is helping us to lower the model size and thereby keep costs in check.</em>”</p><p>In my opinion, Zoho is right in saying that AI models are going to be commoditized. Leveraging AI in a business application is already now a table stake.</p><p>But what about AI and ethics, especially when it comes to training models? After all, Zoho takes a strong position on privacy.</p><p>Ram extends this into the realm of training models. “<em>We strongly believe that your data is your data. And it should work just for you, meaning you cannot use your data to train a third party's AI and in turn, you get some subsidized software. And that's not how it should work. Because that's your business secret, right? That's your secret sauce. And you don't want it to be taken over by some model. So, we have set clear policy privacy guidelines, where we have built foundational models that are independent of any of our customer data. And then they are fine tuned to individual customers.</em>” He insists that “<em>if I'm going to use the model that is trained on one company, for their competitor, who has just signed up for CRM, I'm basically selling their data. We don't do that. So, all of these privacy policies intact, and we also have a strong policy towards making our AI bias free.</em>”</p><p>On top of this, Zoho makes sure that there is no PII used for learning. Drift is controlled via regular systems audits using split systems.&nbsp;</p><p>Zoho promises to do it right, which is, as Ram says, the promise that every vendor should give and keep.&nbsp;</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Mon, 19 Feb 2024 19:41:57 -0500</pubDate></item><item><title><![CDATA[Beyond the hype - How to use chatGPT to create value]]></title><link>https://www.aheadcrm.co.nz/blogs/post/beyond-the-hype-how-to-use-chatgpt-to-create-value</link><description><![CDATA[Now, that we are in the middle of – or hopefully closer to the end of – a general hype that was caused by Open AI’s ChatGPT, it is time to reemphasize ]]></description><content:encoded><![CDATA[<div class="zpcontent-container blogpost-container "><div data-element-id="elm_6DNaBtn1RBun7DtolWbM0w" data-element-type="section" class="zpsection "><style type="text/css"></style><div class="zpcontainer-fluid zpcontainer"><div data-element-id="elm_8OBBKXW6RyWISn7cM7WErg" data-element-type="row" class="zprow zprow-container zpalign-items- zpjustify-content- " data-equal-column=""><style type="text/css"></style><div data-element-id="elm_Ne-NfMhwT1KHVM4CyCI6Dw" data-element-type="column" class="zpelem-col zpcol-12 zpcol-md-12 zpcol-sm-12 zpalign-self- "><style type="text/css"></style><div data-element-id="elm_wnZfsvEDRLOajMYVHhA1fQ" data-element-type="text" class="zpelement zpelem-text "><style></style><div class="zptext zptext-align-center " data-editor="true"><div><p>Now, that we are in the middle of – or hopefully closer to the end of – a general hype that was caused by Open AI’s ChatGPT, it is time to reemphasize on what is possible and what is not, what should be done and what not. It is time to look at business use cases that are beyond the hype and that can be tied to actual business outcomes and business value.</p><p>This, especially, in the light of the probably most expensive demo ever, after<a href="https://www.theverge.com/2023/2/8/23590864/google-ai-chatbot-bard-mistake-error-exoplanet-demo"> Google Bard gave a factually wrong answer</a> in its release demo. A factual error wiped more than $100bn US off Google’s valuation.</p><p>I say this without any gloating. Still, this incident shows how high the stakes are when it comes to large language models, LLM. It also shows that businesses need to have a good and hard look at what problems they can meaningfully solve with their help. This includes quick wins as well as strategic solutions.</p><p>From a business perspective, there are at least two dimensions to look at when assessing the usefulness of solutions that involve large language models, LLM.</p><p>One dimension, of course, is the degree of language fluency the system is capable of. Conversational user interfaces, exposed by chatbots or voice bots and digital assistants, smart speakers, etc. are around for a while now. These systems are able to interpret the written or spoken word, and to respond accordingly. This response is either written/spoken or by initiating the action that was asked for. One of the main limitations of these more traditional conversational AI systems is that they are better in understanding than in – lacking a better word – expressing themselves. Relying on well-trained machine learning models, they are also quite regularly able to surface a correct solution for problems <strong>in the problem domain that they are trained for</strong>. They usually work based on pretrained intents.</p><p>And, based on the training data, they usually give quite accurate responses to questions in their domain.</p><p>The problem: They are usually limited to a fairly small number of domains.</p><p>LLMs, on the other hand, are generally trained “<em>to understand the relationships between words, phrases and sentences in a language. The goal is to have the LLM generate outputs that are semantically meaningful and reflect the context of the input.</em>” This is part of ChatGPTs answer to the question what the purposes of an LLM is. The training set of an LLM is usually a vast amount of “real world” knowledge that usually comes from publicly available sources – aka the Internet. The output itself can be in written, graphical or other formats.</p><p>What LLMs excel in is generating responses to questions in a human way. And they can respond to a wide variety of topics. When focusing on text, they are built to generate coherent and meaningful responses.</p><p>The problem: They sometimes lack accuracy and give wrong output with full confidence. Even worse, wrong or inaccurate output is not easily identifiable by a user without the requisite knowledge. Again, refer to the Google Bard example that (temporarily, at least) wiped off $100 billion US from Googles valuation. Not picking on Google, there are plenty of examples around that call out ChatGPT or You.com or other tools.</p><p>Consequently, the other dimension to look at is accuracy.</p><p>The question is whether both dimensions always matter equally or not. In a business sense, one can argue that accuracy matters always. Receiving factual errors in a business conversation is not only a poor customer experience but may in extreme cases even lead to legal issues.</p><p>What is also important to understand is that the more accuracy is required the more the necessity of integrating additional systems to augment the LLM increases. An LLM on its own is not much more than some form of entertainment. Even in search engines, LLMs only augment the search by enabling natural language queries and the delivery of results in human language instead of a mere link list.</p><p>At least they should do this.</p><p>With all this being said, what are business use cases involving a large language model? As said, there needs to be a reasonable accuracy. Obviously, they require fluency as a precondition, as fluency is the core differentiator of an LLM.</p><p>Let’s look at some use cases in no particular order of priority.</p><figure class="wp-block-image is-resized is-style-default"><a href="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEhYD0c9Q9qhi_kVFzzGY0GgyRyBNRqPzgt_mpPlWmavX8xdM3Evar3Ja-xcb3wAT8iqDhYlb9WXzqWxWSXCxr7wBGPkrfgPcMDs6WRD04-L_2VXINUd8nnc2QymPkIuVN3_TMAULMf44DQVaKqELBRK-oAik9CY7mAJl5i3aa7av-nOZKfqqS-7IMQmZg/s1251/LLM%20Scenarios.png"><img src="https://blogger.googleusercontent.com/img/b/R29vZ2xl/AVvXsEhYD0c9Q9qhi_kVFzzGY0GgyRyBNRqPzgt_mpPlWmavX8xdM3Evar3Ja-xcb3wAT8iqDhYlb9WXzqWxWSXCxr7wBGPkrfgPcMDs6WRD04-L_2VXINUd8nnc2QymPkIuVN3_TMAULMf44DQVaKqELBRK-oAik9CY7mAJl5i3aa7av-nOZKfqqS-7IMQmZg/w640-h360/LLM%20Scenarios.png" alt="LLM business use cases that can be implemented already now" width="837" height="470" title="LLM business use cases that can be implemented already now"/></a><figcaption>LLM business use cases that can be implemented already now</figcaption></figure><ul><li>I’d start with something that I’d call “storytelling”. This is basically the creation of market-relevant documents that describe the capabilities and differentiating factors of a product, solution, or service. Being somewhat marketing related (no offence intended) and a first point of contact for customers, it needs to be easy to understand without requiring a great deal of technical accuracy. At the same time, it must not be wrong. A stripped-down version of this could be the (improved) generation of social media content, e.g., tweets. Benefits are faster creation of high-level content for general websites but also, more specifically, for ABM scenarios and landing pages. To be able to create this text, an LLM needs to be connected to internal systems holding requirements, specifications as well as communications between the involved persons. This is also a use case that should be implement-able near-term.</li><li>One of the main tasks of people is the writing of, and more so, responding to emails. Especially, in sales scenarios, customer inquiries can get formulated and suggested based upon previous emails and the context given by the CRM system, e.g., about proposals made. This scenario would already require quite a high accuracy to avoid sending out faulty information that might be legally binding. The benefit of this scenario is a significant reduction time needed to send emails, resulting in increased productivity. It is a scenario that Microsoft has already implemented in its<a href="https://youtu.be/U5emr9KyquA"> Viva Sales</a> solution.</li><li>Generation of documentation is a scenario that somewhat varies in the requirement for fluency. It can be mainly divided into technical and user documentation. While user documentation needs to be extremely readable, the writing style is somewhat less important for technical documentation. Conversely, technical documentation likely needs to have a high degree of technical accuracy that is not needed in user documentation, which means that either different repositories or different parts of source documents need to be used to create the texts and potentially diagrams and images.</li><li>One of the most promising use cases in the short term is customer service, including enterprise search. Here, users want answers to their questions, not just links or something actioned. To achieve this, it is necessary to connect to a conversational AI, business systems and a well-functioning knowledge base that helps in generating accurate answers when searching for something. The actioning of issues is very similar to what conversational AIs do already now. The differences are that the intent detection can be far better as the LLM can create more than enough training sets for this and that the answers given by the system are far more fluent. The same holds true for an inquiry scenario. However, as a word of caution, the accuracy of responses to inquiries depends heavily on the kb content that gets searched by the enterprise search. Therefore, the kb needs continuous and rigorous scrutiny. If this is given, the benefits lie in increased call deflection and customer satisfaction. Properly implemented, benefits include an improved call deflection as more cases can get handled by the system, combined with an increased customer satisfaction as issue handling can become quite easy and efficient for the customer.<a href="https://www.cognigy.com/"> Cognigy</a> has recently presented some very good examples (<a href="https://youtu.be/WKJO4_JfIFs">here</a> and<a href="https://youtu.be/yZi-0XAZLz0"> here</a>) that also include voice in- and output.</li><li>Agent assistance is somewhat easier to implement, as it mostly needs to connect to the customer service application, including the chat history. Having complete access to sales and marketing data, of course is helpful, too. Combined with a sentiment analysis, the LLM can suggest text blocks for the agent to use. The benefits of this are an increased agent efficiency and quite possibly also higher customer satisfaction as the text blocks do exhibit more empathy with the customer’s situation than texts generated without an LLM.</li></ul><p>In summary, these five scenarios show use cases involving an LLM that are beyond the hype. They can get implemented in a short time and they can also be easily tied to business outcomes. That way, their benefits can get measured.</p><p>Which other use cases do you see? And how would you tie them to business value?&nbsp;</p></div></div>
</div></div></div></div></div></div> ]]></content:encoded><pubDate>Wed, 15 Feb 2023 20:51:11 -0500</pubDate></item></channel></rss>