# TestingCatalog AI News > TestingCatalog covers latest AI news, model releases, leaks, and rumours from the world of artificial intelligence Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### Legal Notice URL: https://www.testingcatalog.com/imprint/ Last updated: 2026-03-13T22:33:16.000Z **Disclaimer** The information contained in this website is for general information purposes only. The information is provided by TestingCatalog and while we endeavour to keep the information up to date and correct, we make no representations or warranties of any kind, express or implied, about the completeness, accuracy, reliability, suitability or availability with respect to the website or the information, products, services, or related graphics contained on the website for any purpose. Any reliance you place on such information is therefore strictly at your own risk. In no event will we be liable for any loss or damage including without limitation, indirect or consequential loss or damage, or any loss or damage whatsoever arising from loss of data or profits arising out of, or in connection with, the use of this website. Through this website, you are able to link to other websites which are not under the control of TestingCatalog. We have no control over the nature, content, and availability of those sites. The inclusion of any links does not necessarily imply a recommendation or endorse the views expressed within them. Every effort is made to keep the website up and running smoothly. However, TestingCatalog takes no responsibility for, and will not be liable for, the website being temporarily unavailable due to technical issues beyond our control. **Imprint** The following information (Impressum) is required under German law. - Responsible for content (§ 5 DDG / § 18 (2) MStV): Alexey Shabanov, Talstrasse 3, 13189, Berlin, Germany - Contact information: testingcataloghelp@gmail.com - Management: Alexey Shabanov - VAT No. in accordance with § 27 a VAT Law: Not applicable - Liability: Despite diligent control over content, we do not assume any liability for the content of external links. The content of external websites is the exclusive responsibility of their respective operators. ### Terms and Conditions ("Terms") URL: https://www.testingcatalog.com/terms/ Last updated: 2026-03-13T22:22:55.000Z Last updated: March 09, 2026 Please read these Terms and Conditions ("Terms", "Terms and Conditions") carefully before using the [testingcatalog.com](https://www.testingcatalog.com/) website (the "Service") operated by TestingCatalog ("us", "we", or "our"). This agreement sets forth the legally binding terms and conditions for your use of the site at URL. These Terms apply to all visitors, users and others who access or use the Service. By accessing or using the site in any manner, including, but not limited to, visiting or browsing the site or contributing content or other materials to the site, you agree to be bound by these terms and conditions. If you disagree with any part of the terms then you may not access the Service. **Security** TestingCatalog has sophisticated security measures in place for URL to help protect against the loss, misuse, and alteration of the data under TestingCatalog’s control. When accessing TestingCatalog via a supported Web Browser, Secure Socket Layer (SSL) technology protects information using both server authentication and data encryption to help ensure data is safe, secure, and available only to you. TestingCatalog also implements an advanced security method based on dynamic data and encoded session identifications and hosts URL in a secure server. Finally, TestingCatalog requires unique user names and passwords that must be entered each time a member logs on. These safeguards help prevent unauthorized access, maintain data accuracy, and ensure the appropriate use of data. **Accounts** When you create an account with us, you must provide us information that is accurate, complete, and current at all times. Failure to do so constitutes a breach of the Terms, which may result in immediate termination of your account on our Service. We require our members to provide personal contact information, including name, address and e-mail address. We also ask for personal information such as age, gender, and other data. You are responsible for safeguarding the password that you use to access the Service and for any activities or actions under your password, whether your password is with our Service or a third-party service. You agree not to disclose your password to any third party. You must notify us immediately upon becoming aware of any breach of security or unauthorized use of your account. **Links To Other Web Sites** Our Service may contain links to third-party websites or services that are not owned or controlled by testingcatalog. TestingCatalog has no control over and assumes no responsibility for, the content, privacy policies, or practices of any third party websites or services. You further acknowledge and agree that TestingCatalog shall not be responsible or liable, directly or indirectly, for any damage or loss caused or alleged to be caused by or in connection with the use of or reliance on any such content, goods or services available on or through any such websites or services. We strongly advise you to read the terms and conditions and privacy policies of any third-party websites or services that you visit. **Termination** We may terminate or suspend access to our Service immediately, without prior notice or liability, for any reason whatsoever, including without limitation if you breach the Terms. All provisions of the Terms which by their nature should survive termination shall survive termination, including, without limitation, ownership provisions, warranty disclaimers, indemnity and limitations of liability. We may terminate or suspend your account immediately, without prior notice or liability, for any reason whatsoever, including without limitation if you breach the Terms. Upon termination, your right to use the Service will immediately cease. If you wish to terminate your account, you may simply discontinue using the Service. All provisions of the Terms which by their nature should survive termination shall survive termination, including, without limitation, ownership provisions, warranty disclaimers, indemnity and limitations of liability. **Content** Our Service allows you to post, link, store, share and otherwise make available certain information, text, graphics, videos, or other material ("Content"). You are responsible for the content, that you submitted. It is allowed to post Bug Reports, Test Cases, User Stories, User Feedbacks and other information related to the Test Session, including screenshots, videos, links and log files only if Beta program is opened for everyone (public). Otherwise, you should act accordingly to the Terms and Conditions of the third party service of related Project. It is not allowed to submit any information about security issues. It is not allowed to post any harmful or illegal information. **Governing Law** These Terms shall be governed and construed in accordance with the laws of Berlin, Germany, without regard to its conflict of law provisions. Our failure to enforce any right or provision of these Terms will not be considered a waiver of those rights. If any provision of these Terms is held to be invalid or unenforceable by a court, the remaining provisions of these Terms will remain in effect. These Terms constitute the entire agreement between us regarding our Service, and supersede and replace any prior agreements we might have between us regarding the Service. **Changes** We reserve the right, at our sole discretion, to modify or replace these Terms at any time. If a revision is material we will try to provide at least 30 days notice prior to any new terms taking effect. What constitutes a material change will be determined at our sole discretion. By continuing to access or use our Service after those revisions become effective, you agree to be bound by the revised terms. If you do not agree to the new terms, please stop using the Service. **Contacts** If you have any questions about these Terms, please contact us. ### Privacy Policy ("Policy") URL: https://www.testingcatalog.com/policy/ Last updated: 2026-03-13T22:59:24.000Z Last updated: March 09, 2026 Your privacy is very important to us. We protect it from third parties and from other users of this service. **Information you share with us** We need to collect information about you to provide you with the Services or the support you request. Additionally, you can choose to voluntarily provide information to us. Information we collect from you Account information: If you choose to create a TestingCatalog account, you must provide us with some personal information, such as your username, email address, a password, city, country, your expertise and your primary role as a beta tester. Following the initial account creation, you may choose to complete your profile by providing additional information such as your birth date (displayed only as age in years), a secondary role as developer, interest in apps, time availability, experience with software testing and/or Android development, links to your website or your publicly accessible social profile, your name and whether you are seeking beta testers for your own Android app or to become a beta tester. Some of our product features, such as searching and viewing public TestingCatalog beta tester profiles do not require you to create an account. On TestingCatalog, your profile including your username is always listed publicly. You can use either your real name or a pseudonym as your username. Your name and email address are kept private from other members of the service and used for communications between TestingCatalog and you. Social media account information: You have to sign in to your TestingCatalog account with your Google credentials. If you choose to do this, the first time you do so you will be asked whether you agree that Google may provide certain information to TestingCatalog, such as your name, email address, profile photo, city, birthdate, profile link, professional experience, expertise, interests and other information associated with your social media account. We use this information to help you create your account and complete your profile without having to type this information in manually. Communication information: You may choose to visit other members’ TestingCatalog profiles and communicate with them by living comments on their public profile pages. When you do so, we process information about which profiles you visit and store messages you send to other users. App information: You may choose to post a link to your app and share it with others. Just like profiles, apps are public information on TestingCatalog. When you post an app you may share information including a title, description, tags, stage of development, the type of help you’re seeking, and links to further information, e.g. a website or video. Upgrade information: You may choose to upgrade your TestingCatalog account with a paid Pro Plan (In development). When you do so, you provide us with additional information including your billing address, and optionally a company name and VAT number. We record this information together with what plan you purchased and when. We also display a Pro member badge next to your profile to help you stand out from the crowd and find beta testers faster. Other information: Information that you voluntarily provide to us, including your survey responses, participation in contests, promotions, suggestions for improvements, or any other actions performed on the Services. Information we collect from your use of our Services We collect information about you and the devices you use to access the Services, such as your computer, mobile phone, or tablet. The information that we collect includes: Device Information: Information about your device, including your hardware model, operating system and version, device name, unique device identifier, mobile network information, and information about the device’s interaction with our Services. Use Information: Information about how you use our Services, including your access time, “login” and “logout” information, browser type and language, country and language setting on your device, the domain name of your Internet service provider, other attributes about your browser, mobile device and operating system, any specific page you visit on our platform, content you view, features you use, the date and time of your visit to or use of the Services, your search terms, the website you visited before you visited or used the Services, data about how you interact with our Services, and other clickstream data. **How we use your information** We use your information primarily to help you find beta testers or apps for beta testing. And we use it secondarily to improve our Services to better help you find beta testers or apps for beta testing. In order to do so, we may use information about you for a number of purposes, including: Providing, improving, and developing our Services - Processing or recording payment transactions or money transfers - Otherwise providing you with the products and features you choose to use - Providing, maintaining and improving our Services - Developing new products, services and improvements - Delivering the information and support you request, security alerts, support and administrative messages, and provide assistance for problems with our Services - Improving, personalizing, and facilitating your use of our Services - Measuring, tracking, and analyzing trends and usage in connection with your use or the performance of our Services Communicating with you about our Services - An integral part of TestingCatalog are the communications we send to help you find beta testers. These include notifications, for example, when someone comments on an app you shared, visits your profile, or sends you a comment, as well as our weekly newsletter, which includes useful information regarding everything Android apps and software testing related. Should you choose to no longer receive the weekly newsletter or other notifications, you can do so any time either by adjusting your message settings, or clicking the unsubscribe link contained in the email’s footer. Protecting our services and maintaining a trusted environment - Investigating, detecting, preventing, or reporting fraud, misrepresentations, security breaches or incidents, other potentially prohibited or illegal activities, or to otherwise help protect your account - Protecting the security or integrity of our Services - Enforcing our Terms of Service or other applicable agreements or policies - Complying with any applicable laws or regulations, or in response to lawful requests for information from the government or through legal process - Fulfilling any other purpose disclosed to you in connection with our Services - Contacting you to provide assistance with our Services Advertising and marketing - Marketing of our Services - Communicating with you about opportunities, products, services, contests, promotions, discounts, incentives, surveys, and rewards offered by us and select partners - If we send you marketing emails, each email will contain instructions permitting you to opt out of receiving future marketing or other communications Other uses - For any other purpose disclosed to you in connection with our Services from time to time **Information we share & disclose** We may share information about you as follows: With other users of our Services with whom you interact - With other users of our Services with whom you interact through your own use of our Services. For example, we may share information when you visit another user’s profile or contact them by leaving a comment, or when they visit your publicly available profile. With third parties - With third parties to provide, maintain, and improve our Services, including service providers who access information about you to perform services on our behalf (e.g., fraud prevention, identity verification, and fee collection services), as well as financial institutions, payment networks, payment card associations, credit bureaus, partners providing services on TestingCatalog’s behalf, and other entities in connection with the Services. - With third parties that have joined our partner network and for who we power beta tester discovery on their websites. - With third parties that run advertising campaigns, contests, special offers, or other events or activities on our behalf or in connection with our Services. Business transfers and corporate changes - To a subsequent owner, co-owner, or operator of our Services or - In connection with (including, without limitation, during the negotiation or due diligence process of) a corporate merger, consolidation, or restructuring; the sale of substantially all of our stock and/or assets; financing, acquisition, divestiture, or dissolution of all or a portion of our business; or other corporate change. Safety and compliance with law - If we believe that disclosure is reasonably necessary (i) to comply with any applicable law, regulation, legal process or governmental request (e.g., from tax authorities, law enforcement agencies, etc.); (ii) to enforce or comply with our Terms of Service or other applicable agreements or policies; (iii) to protect our or our customers’ rights, or the security or integrity of our Services; or (iv) to protect us, users of our Services from harm, fraud, or potentially prohibited or illegal activities. With your consent - When you authorize a third party application or website to access your information, for example by sharing your profile or app on social media. Aggregated and anonymized information - We also may share aggregated and anonymized information that does not specifically identify you or any individual user of our Services. **Cookies and similar technologies** We use various technologies to collect information when you access or use our Services, including placing a piece of code, commonly referred to as a “cookie,” or similar technology on your device and using web beacons. Cookies are small data files that are placed on your computer or mobile device when you visit a website. Cookies are widely used by website owners in order to make their websites work, or to work more efficiently, as well as to provide reporting information. We use cookies to better understand how people use our Service. For example, cookies can help us understand how people use TestingCatalog, analyze which parts of the Service people find most useful and engaging, and identify features that could be improved. Cookies set by the website owner (in this case, TestingCatalog) are called “first party cookies”. Cookies set by parties other than the website owner are called “third party cookies”. Third party cookies enable third party features or functionality to be provided on or through the website (e.g. analytics, interactive content and advertising). The parties that set these third party cookies can recognise your computer both when it visits the website in question and also when it visits certain other websites. Why do we use cookies? Some cookies are required for technical reasons in order for our Service to operate. Other cookies also enable us to track and target the interests of our users to enhance the experience of our Services. Third parties serve cookies through our websites for analytics and other purposes. We use different types of cookies: Essential website cookies: These cookies are strictly necessary to provide you with the Services available through our websites and to use some of its features, such as access to secure areas. Because these cookies are strictly necessary to deliver our Service to you, you cannot refuse them without impacting how our Service functions. You can block or delete them by changing your browser settings however, as described below under “How can I control cookies?” Performance and functionality cookies: These cookies are used to enhance the performance and functionality of our Services but are non-essential to their use. However, without these cookies, certain functionality may become unavailable. To refuse these cookies, please follow the instructions below under “How can I control cookies?” Analytics and customisation cookies: These cookies collect information that is used either in aggregate form to help us understand how our websites are being used or how effective our marketing campaigns are, or to help us customise our Services for you in order to enhance your experience. These cookies enable analytics reporting for visitor engagement on the websites. To refuse these cookies, please follow the instructions below under “How can I control cookies?” How can I control cookies? Most web and mobile device browsers are set to automatically accept cookies by default. However, you have the right to decide whether to accept or reject cookies. You can set or amend your web browser controls to accept or refuse cookies, as well block or delete them. If you choose to reject cookies, you may still use our Service though your access to some functionality and areas of our websites may be restricted. As the means by which you can refuse cookies through your web browser controls vary from browser-to-browser, you should visit your browser’s help menu for more information. Alternatively, to adjust your Facebook cookie preferences, please visit https://www.facebook.com/policies/cookies. To opt out of Google or DoubleClick cookies please visit http://preferences-mgr.truste.com/. You also can learn more about cookies by visiting http://www.allaboutcookies.org, which includes additional useful information on cookies and how to block cookies on different types of browsers and mobile devices. Please note, however, that by blocking or deleting cookies used in the Services, you may not be able to take full advantage of the Services. We also may collect information using web beacons. Web beacons are electronic images that may be used in our Services or emails. We use web beacons to deliver cookies, track the number of visits to our website, understand usage and campaign effectiveness, and determine whether an email has been opened and acted upon. **Third party services and analytics** We use third-party service providers to provide site metrics and other analytics services. These third parties can use cookies, web beacons, and other technologies to collect information, such as your IP address, identifiers associated with your device, other applications on your device, the browsers you use to access our Services, web pages viewed, time spent on webpages, links clicked, and conversion information (e.g., transactions entered into). This information can be used by us and third-party service providers on our behalf to analyze and track usage of our Services, determine the popularity of certain content, and better understand how you use our Services. The third-party service providers that we engage are bound by confidentiality obligations and other restrictions with respect to their use and collection of your information. This Privacy Policy does not apply to, and we are not responsible for, third-party cookies, web beacons, or other tracking technologies, which are covered by such third parties’ privacy policies. For more information, we encourage you to check the privacy policies of these third parties to learn about their privacy practices. Examples of our third-party service providers to help deliver our Services or to connect to our Services include: - Google Analytics: We use Google Analytics to understand how our Services perform and how you use them. To learn more about how Google processes your data, please visit https://www.google.com/policies/privacy/. To opt out of Google Analytics please visit https://tools.google.com/dlpage/gaoptout. - Google Adsense: https://www.google.com/adsense - OneSignal: https://onesignal.com These third party service providers make use of cookies to implement their services. **Children** Our Services are general audience services not directed at children, and you may not use our Services if you are under the age of 13\. You must also be old enough to consent to the processing of your personal data in your country (in some countries we may allow your parent or guardian to do so on your behalf). If we obtain actual knowledge that any information we collect has been provided by a child under the age of 13, we will take steps delete the information as soon as possible. **Managing your information with us** We generally retain your information as long as reasonably necessary to provide you the Services or to comply with applicable law. If you reside in the European Union, you have the right under the General Data Protection Regulation to request from TestingCatalog access to and rectification or erasure of your personal data, data portability, restriction of processing of your personal data, the right to object to processing of your personal data, and the right to lodge a complaint with a supervisory authority. If you reside outside of the European Union, you may have similar rights under your local laws. To request access to or rectification, portability or erasure of your personal data, or to delete your TestingCatalog account, please see below. Personal information You may access, change, or correct information that you have provided by logging into your account and update your profile at any time. You may request a copy of your personal information by submitting a request to our customer support. Deactivating or deleting your account If you wish to deactivate or delete your account, you can do so by contacting us here. Promotional communications You can opt in and out of receiving promotional messages from TestingCatalog by following the instructions in those messages, or by changing your notification settings by logging into your TestingCatalog account. Opting out of receiving communications may impact your use of the Services. If you decide to opt out, we can still send you non-promotional communications about your account or our ongoing business relations. **Security** Your account is protected by the Google account you choose. Just as with most Internet services, the security of your information depends on this password, and how well you safeguard access to your account from others. Fortunately, this is not very difficult. You can prevent unauthorized access to your account by choosing a longer password containing numbers and letters, by limiting the access to your computer and browser, and by logging out from your account when you are done using the service. Should you have accidentally disclosed your password to someone, or otherwise come to suspect someone might be accessing your account, you can easily change your password by choosing a new one in your profile information. While we would certainly like to be able to, we cannot guarantee the security of user information at all times. We take reasonable measures, including administrative, technical, and physical safeguards, to protect your information from loss, theft, misuse, and unauthorized access, disclosure, alteration, and destruction. Nevertheless, the internet is not a 100% secure environment, and we cannot guarantee absolute security of the transmission or storage of your information. We hold information about you both at our own premises and with the assistance of third-party service providers. **Storage and processing** We may, and we may use third-party service providers to, process and store your information in the European Union, United States, and other countries. Changes to this Privacy Policy We may amend this Privacy Policy from time to time by posting a revised version and updating the “Effective Date”. The revised version will be effective on the “Effective Date” listed. We will provide you with reasonable prior notice of material changes in how we use your information, including by email, if you have provided an email address. If you disagree with these changes, you may cancel your account at any time. Your continued use of our Services constitutes your consent to any amendment of this Privacy Policy. **Contact** You may contact us with any questions or concerns regarding this Privacy Policy by submitting a request to our support. If you have any questions or concerns regarding this policy, or if you believe our policy or applicable laws relating to the protection of your personal information have not been respected, you may file a complaint, and we will respond to let you know who will be handling your matter and when you can expect a further response. We may request additional details from you regarding your concerns and may need to engage or consult with other parties in order to investigate and address your issue. We may keep records of your request and any resolution. Effective Date: March 09, 2026 ### About us URL: https://www.testingcatalog.com/about/ Last updated: 2026-05-10T21:00:00.000Z > **TestingCatalog’s only official website is testingcatalog.com.** > TestingCatalog is not affiliated with testingcatalog\[.\]net or any similarly named domains. The official TestingCatalog publication is founded and operated by Alexey Shabanov from Berlin, Germany. ## What is TestingCatalog? TestingCatalog is an AI feature intelligence publication covering the latest changes in consumer and developer AI products. We track announced features, unreleased features, model launches, interface experiments, product rollouts, and early signals from companies building the next generation of AI tools. Our readers include early adopters, builders, researchers, product teams, journalists, investors, and anyone who wants to understand what is changing in AI products before it becomes obvious. ## What we cover TestingCatalog focuses on practical AI product changes, including: - new and upcoming features in ChatGPT, Gemini, Claude, Grok, Perplexity, Copilot, Mistral, and other AI products - model launches, model upgrades, and product integrations - mobile, web, API, and developer-tool changes - limited rollouts, A/B tests, regional availability, and feature flags - AI agents, coding tools, search products, productivity tools, and creative AI workflows - confirmed announcements, observed tests, leaks, rumors, and industry signals ## How we report TestingCatalog’s reporting starts with testing. We monitor live products, public builds, release notes, developer documentation, official announcements, visible interface changes, public product assets, and community-discovered signals. When possible, we test features directly and explain what changed, who can access them, and why they matter. **We separate evidence from interpretation:** - \- Confirmed, means the information comes from official announcements or public documentation. - \- Rolling out, means the feature is visible to some users, platforms, or regions, but not everyone. - \- Testing, means the feature appears in experiments, has limited access or exhibits early product behavior. - \- Rumor / Speculation, means the information is not officially confirmed and should be treated as provisional. - \- Analysis, means TestingCatalog’s interpretation of available signals. ## Editorial standards TestingCatalog aims to make fast-moving AI product news useful, sourced, and understandable. **Our editorial principles:** - \- Test features directly whenever possible - \- Link to original sources when available - \- Distinguish official announcements from observed tests and rumors - \- Correct or clarify inaccurate information - \- Disclose sponsored posts and partnerships - \- Use AI tools for editing and production support while keeping editorial responsibility with TestingCatalog ## Our History TestingCatalog started in 2014 as a beta-testing community for Android app developers and early adopters. In 2016, TestingCatalog.com launched as a catalog for beta Android apps. In 2018, TestingCatalog evolved into a news platform covering unreleased app features, later expanding to broader coverage of technology and AI. Today, TestingCatalog focuses on AI feature intelligence: what is launching, what is being tested, what is changing, and what those changes mean for users and builders. TestingCatalog reporting has been referenced by publications including TechCrunch, Bloomberg, Android Authority, Android Police, XDA, SocialMediaToday, and others. ## Our team TestingCatalog works with contributors, researchers, editors, and AI-assisted workflows to monitor product changes across the AI ecosystem. Each author page includes contributor background, coverage areas, and published posts. ## Contact, corrections & tips For news tips, corrections, partnerships, advertising, or official inquiries, contact TestingCatalog through the official Contact page or the email listed on testingcatalog.com. If you see another website or account claiming to represent TestingCatalog, please verify it against this page. The official TestingCatalog website is testingcatalog.com. ### Telegram Beta URL: https://www.testingcatalog.com/telegram-beta/ Last updated: 2023-12-04T20:53:16.000Z Telegram is the only messaging application on Android offering a simple and easy-to-use, quick user experience, with full end-to-end encryption, unlimited storage, and no pesky ads. Apart from one-on-one and group messaging, you can create your channels and broadcast content to your followers. ## How to become a Telegram beta tester? The beta version of the Telegram app is not distributed via Google Play, unlike Telegram X. Telegram Beta has a different package id and can be installed as a standalone app alongside the stable version. To get it, you need to install the APK file by yourself. The easiest way is to visit our [@tgtester](https://telegram.me/tgtester?ref=testingcatalog.com) channel on Telegram and grab the latest available version. Further updates will be distributed as automatic in-app updates. You can find more details on that from our ["How to become a beta tester for Telegram"](https://www.testingcatalog.com/how-to-join-telegram-beta-and-telegram-x-beta-on-android/) post. [How to become a beta tester for Telegram apps on Android (Telegram, Telegram X and Plus Messenger)Updated on March 2020\. Currently, there are three official Telegram clients for Android - Telegram, Telegram X and Plus Messenger. All of them have a beta program but they are distributed in different ways.![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2020/03/Adobe_Post_20200322_1725350.6811167077944059--1-.jpg)](https://www.testingcatalog.com/how-to-join-telegram-beta-and-telegram-x-beta-on-android/) ### How to download the official Telegram beta build 1. Head over to the [App Center page](https://install.appcenter.ms/users/drklo-2kb-ghpo/apps/telegram-beta-2/distribution%5Fgroups/all-users-of-telegram-beta-2?ref=testingcatalog.com) for Telegram beta. 2. Download the latest APK file. 3. Make sure that installation from unknown sources is allowed. 4. Install Telegram beta APK. ## How to leave Telegram beta? Because Telegram beta comes as a standalone APK, there is no special process of leaving it. You can simply install a Stable version from Google Play and uninstall Telegram beta app. ## Side-load Telegram APKs Telegram beta APK can be also downloaded from [unofficial Telegram channels](https://telegram.me/tgtester?ref=testingcatalog.com) or [APKMirror](https://www.apkmirror.com/apk/telegram-fz-llc/telegram/?ref=testingcatalog.com). The installation and update process will remain the same. ## Updating Telegram beta Telegram beta has in-app updates process build-in. It means that if you have a beta APK installed, as soon as a new update will be released, you will see a pop-up in the app asking you if you want to download a new update. This means that you will need to side-load beta APK only once and Telegram will handle all further updates for you. ## Telegram stable releases Telegram beta normally receives new versions months or weeks ahead of its stable version. Once a major beta update is out, you will receive a bunch of minor bug fix update afterwards. After beta testing is finished, a new major version will be released on Google Play along with a blog post that explains all newly released features in detail. Until a new major beta release, beta and stable apps will remain identical for some time. [Telegram - Apps on Google PlayTelegram is a messaging app with a focus on speed and security.![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayTelegram FZ-LLC![](https://play-lh.googleusercontent.com/pgiel6OzR01do2sPxONlteJ5_GAwD6W-JpiF98hdns3XinShKZkziUd1B8sG-8qa)](https://play.google.com/store/apps/details?id=org.telegram.messenger&ref=testingcatalog.com) Besides the stable app that is distributed over Google Play, Telegram has another standalone stable version that is distributed [directly from their website](https://telegram.org/android?ref=testingcatalog.com). This app is supposed to be almost the same as stable but with fewer restrictions over its content and with a bit more frequent updates. ## Telegram channels and bots Telegram has several official channels where can find some information about updates and recent changes. However, there is no official channel about Telegram beta and that's why @tgtester exists, to cover Telegram updates before they are getting released to everyone. ### Unofficial Telegram channels - [@tgtester](https://telegram.me/tgtester?ref=testingcatalog.com) \- Unofficial channel by TestingCatalog about Telegram beta apps for Android. ### Official Telegram channels - [Telegram news](https://telegram.me/telegram?ref=testingcatalog.com) \- Official channel about Telegram news and stable updates. - [Durov channel](https://telegram.me/durov?ref=testingcatalog.com) \- A personal channel of Pavel Durov, Telegram founder. - [Telegram tips](https://telegram.me/telegramtips?ref=testingcatalog.com) \- A channel about stable features of Telegram and also tips and tricks. - [Memes channel](https://telegram.me/MemesTelegram?ref=testingcatalog.com) \- Telegram memes channel that was used in the past to stress test some new Telegram features. For example, it was used to test channel voice chats while this feature was in beta. - [Telegram contests](https://telegram.me/contest?ref=testingcatalog.com) \- A channel for developer contests and announcements. Sometimes you can predict which features may come to Telegram in the future by following contests organised by the Telegram team. - [Telegram Public Testing](https://t.me/joinchat/1GTwd3c9OjIzNTEy?ref=testingcatalog.com) \- A relatively new channel that is focused on public stress testing of certain Telegram features. - [Telegram Themes channel](https://telegram.me/AndroidThemes?ref=testingcatalog.com) \- A channel about theming and Telegram customization. ### Useful bots - [Donation bot](https://telegram.me/donate?ref=testingcatalog.com) \- A verified third-party solution that allows users to collect donations on Telegram. - [Controller bot](https://telegram.me/controllerbot?ref=testingcatalog.com) \- A helpful bot for channel admins that can add reactions to your posts, schedule publishing and more. - [Group Help bot](https://telegram.me/grouphelpbot?ref=testingcatalog.com) \- A group moderation bot with a variety of spam protection and other group management features. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/08/781e1a53-691d-4683-bf37-eb12fff2610e.jpeg) Group Help bot on Telegram ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/08/e0c2a0cd-a714-462f-901b-1092786ed11c.jpeg) Donate bot on Telegram Do you know more? Send us a tip on the [Telegram Testers group](https://telegram.me/tgtesterteam?ref=testingcatalog.com) to add it to the list. ## Features Overview ### Secure messenger All Telegram messages are always securely encrypted. Messages in Secret Chats use client-client encryption, while Cloud Chats use client-server/server-client encryption and are stored encrypted in the Telegram Cloud. This enables your cloud messages to be both secure and immediately accessible from any of your devices – even if you lose your device altogether. ### Channels and groups These two are the basic content distribution channels of Telegram. Channels broadcast messages from one user to many and groups allow many users to chat openly. Users can post text, images, videos or any other files to both of them. It makes it a great platform for news brands and application developers who want to distribute their apps without uploading them to Google Play. ### Voice and video chats Voice chats are already available for both, channels and groups while video chats will catch them up at some point. Voice chats became a quite popular format after Clubhouse debut and Telegram is not getting behind. ### Customization and theming Telegram has an advanced customization and theming engine. There you can create a fully custom theme and share it with others. It also has a [verified themes channel](https://telegram.me/AndroidThemes?ref=testingcatalog.com) with tons of different variants to try. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/08/354b436f-e9b0-4cf8-848c-a7d16d6b184c.jpeg) Telegram Stickers and Theming settings ### Open-source nature The client-side of the Telegram app is open-source. It means that other developers can use it in order to create alternative Telegram clients as well. At the same time, it can be audited by any security researcher to make sure that there are no issues with encryption. ### Music and audiobooks player Apart from being a good messenger, Telegram can also work as a music or audiobooks player. Its built-in music player and access to a wide range of different music channels make it a good alternative to full-featured music players. ## Reporting bugs for Telegram Telegram opened a ["Bugs and Suggestions"](https://bugs.telegram.org/?ref=testingcatalog.com) page quite recently where users can report their issues, suggest new features and vote for others. Fixed or implemented features will be marked as Done afterward. Users can also subscribe to conversations and sort items by date or by rating. ## Alternative Telegram clients It is worth mentioning that Telegram also has an official client that is called Telegram X. Apart from this, one of the most popular third-party clients is called Plus Messenger. [Telegram X - Apps on Google PlayInstant messaging — simple, fast, secure, and synced across all your devices.![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayTelegram LLC![](https://play-lh.googleusercontent.com/4Y4oWf0IVCj-nqtqZNpXnCtK-wi5I2XOkGEdA43knOwelyEL_g4KCbtcNjHzFIEXNAsm)](https://play.google.com/store/apps/details?id=org.thunderdog.challegram&ref=testingcatalog.com) [Plus Messenger - Apps on Google PlayPlus Messenger is an UNOFFICIAL messaging app that uses Telegram’s API![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google Playrafalense![](https://play-lh.googleusercontent.com/htMII4ddv4tfyw5omDGtttEmJBRdvjZtGsCvq_DvtM-FIzTVqWeEnQb8uMzStKNBe8P3)](https://play.google.com/store/apps/details?id=org.telegram.plus&ref=testingcatalog.com) ## Links and updates - [Telegram on Google Play](https://play.google.com/store/apps/details?id=org.telegram.messenger&ref=testingcatalog.com) - [Telegram Beta OPT-IN page](https://install.appcenter.ms/users/drklo-2kb-ghpo/apps/telegram-beta-2/distribution%5Fgroups/all-users-of-telegram-beta-2?ref=testingcatalog.com) - [Telegram on APKMirror](https://www.apkmirror.com/apk/telegram-fz-llc/telegram/?ref=testingcatalog.com) If you want to catch up with the latest news about Telegram and other Telegram apps, you will need to [subscribe to our weekly newsletter](https://www.getrevue.co/profile/testingcatalog?ref=testingcatalog.com) with a full summary of different news and updates on these topics 📩 You can also find all recent Telegram features, news and updates that we've reported by navigating to the Telegram tag page 👇 [Telegram - TestingCatalogA news publication about Android for beta testers and early adopters who are eager about trying new apps and features![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov](https://www.testingcatalog.com/tag/telegram/) ## Other Telegram guides [How to enable comments on Telegram channelTelegram allows channel owners to enable native comments functionality by linking a discussion group to the channel![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2020/10/Adobe_Post_20201004_2209340.2676116103224152.png)](https://www.testingcatalog.com/how-to-enable-native-comments-in-telegram/) [How to schedule a voice chat on the Telegram channelTelegram beta v7.7.0 now allows to schedule voice chats to a later time. After the previous release v7.6.0, where Telegram got a voice chat functionality for channels, Telegram continued enhancing this feature even further. Now you can schedule voice chat for a later time and it![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/04/photo5444905336390660484.jpg)](https://www.testingcatalog.com/telegram-beta-v7-7-0-now-allows-to-schedule-voice-chats-to-a-later-time/) [The music player in Telegram - everything you need to knowTelegram is a messaging app in its core, but it also packs some extra features you may haven’t come across. One of them, in particular, is the music player, which can be used to play music files attached to your messages. The player can also play in an order all![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/04/photo5418024845481980768.jpg)](https://www.testingcatalog.com/the-music-player-in-telegram-everything-you-need-to-know/) [How to access debug menu in Telegram beta and enable chat bubblesTelegram beta has a secret debug menu and this menu opens a list of different options for testers and developers. In other words, it exposes an interface for some debugging features. With the recent beta v6.3.0, Telegram devs added an interesting option to this menu that allows you![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2020/07/Adobe_Post_20200722_1529130.4141762212906125.png)](https://www.testingcatalog.com/how-to-access-debug-menu-in-telegram-beta-and-enable-chat-bubbles/) ### Snapchat Beta URL: https://www.testingcatalog.com/snapchat-beta/ Last updated: 2023-12-04T20:53:04.000Z Snapchat is a social media platform, mainly focused on the younger generations and is known best for its AR Lens feature. The app became very popular rapidly and grew its feature set over the past years. ## How to become a Snapchat beta tester? Snapchat for Android has two release tracks on Google Play that are publicly available - stable and beta. You can become a beta tester for Snapchat from both, web and Android Google Play clients. In addition, Snapchat has an option to opt-in for beta testing straight from the app itself. ### From the Snapchat app Previously, it was possible to get a link to the opt-in form on Google Play straight from the settings menu of the Snapchat application. This option is not currently available but it may return in the future. 1. ~~Open the Snapchat app and go to your profile.~~ 2. ~~Hit the gear icon on the top right to open the settings.~~ 3. ~~Scroll down until you find the 'Snapchat Beta' option, select it and tap on the 'Opt-in' to enter the beta program.~~ ### Join on the web 1. Open [Snapchat Beta OPT-IN](https://play.google.com/apps/testing/com.snapchat.android?authuser=1&hl=en&ref=testingcatalog.com) link. 2. Press on the "Become a Tester" button. 3. Press on the ["download it on Google Play"](https://play.google.com/store/apps/details?id=com.snapchat.android&ref=testingcatalog.com) link. 4. See the app title to be updated with "Beta" and update it to the latest version. ### Join on Android 1. Search for Snapchat. 2. Open the Snapchat application page. 3. Scroll down until the "Join the beta" section. 4. Tap on the "Join" button and confirm it on the pop-up window. 5. Wait until the app title to be updated with "Beta" and update it to the latest version. 6. Bear in mind that the OPT-IN process via Google Play may take up to several hours. ![Snapchat beta for Android](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/06/photo5343537841847776823.jpg) Snapchat beta for Android ## What does it mean "beta is full"? If you are trying to join via Google Play for Android, you can try to do the same [via the web link.](https://play.google.com/apps/testing/com.snapchat.android?authuser=1&hl=en&ref=testingcatalog.com) Google Play has a bug that shows "beta is full" status incorrectly and using a web OPT-IN link may be a workaround. However, the "Beta is full" message means that the current testers limit was reached already. If you see this message on the web page, it means that there is no way to join it until Snapchat Devs increase the limit of participants. 📲 Please check the current status on the [official status page for the Snapchat beta](https://www.snapchat.com/beta?ref=testingcatalog.com). ## Why don't I have all features that exist for others? Snapchat for Android has different ways to deliver new features to end-users: - **Client-side updates** \- come as a part of a new APK version. New client-side features usually available to everyone but they are rare. - **Server-side A/B tests** \- these updates make certain features available to a limited amount of users. - **Gradual rollouts** \- a process where Snapchat releases new features to a wider percentage of users. A/B tests and gradual rollouts are the reasons why you may see some features that others don't and another way around. While for gradual rollouts, you may just need to wait for a little bit longer to receive it, with A/B tests it may happen that you won't receive a certain feature at all. ## How to leave the Snapchat beta program? You can leave the Snapchat beta program at any time in the same way as you signed up via Google Play on the web or Android. Leaving the beta program will affect your access to the available set of client-side and server-side features. ### Leave on the web 1. Open [OPT-IN link.](https://play.google.com/apps/testing/com.snapchat.android?authuser=1&hl=en&ref=testingcatalog.com) 2. Press on the "Leave the program" button. 3. Press on the ["download it on Google Play"](https://play.google.com/store/apps/details?id=com.snapchat.android&ref=testingcatalog.com) link. 4. Uninstall the app and reinstall it again from Google Play. ### Leave on Android 1. Search for Snapchat. 2. Open the Snapchat application page. 3. Scroll down until the "You are a beta tester" section. 4. Tap on the "Leave" button and confirm it on the pop-up window. 5. Uninstall the app and reinstall it again from Google Play. Leaving the beta release track means that you will only be receiving stable Snapchat updates. These updates are being released less frequently but if you face any issues with your Snapchat app it could be a way to solve them. Leaving beta release track is recommended in case if you experience technical issues with the app like application crashes or UI freezes for example. ## Side-load from APKMirror It is also possible to side-load Snapchat beta APKs if you want to avoid using Google Play for some reason. The easiest and safest way to do so will be checking and downloading the latest updates from [APKMirror.](https://www.apkmirror.com/?s=com.snapchat.android&post%5Ftype=app%5Frelease&searchtype=apk&ref=testingcatalog.com) There you can find both stable and beta releases with the "beta" prefix in the version name. ## Snapchat Alpha testing Previously, Snapchat was opening its Alpha release track to the public while Snap devs were testing the new UI redesign that you are probably using right now as well. Initially, during the time when this test was internal, it was spotted by reverse engineers and owners of rooted Android devices could get access to this feature by modifying one of its application flags. Later in time, the Snapchat team made their alpha release track available to everyone who knows the trick of how to enable it. In order to do so, users had to navigate to Bermuda island on the SnapMap. When this Alpha test became promoted to the beta release track, It became closed again. ## Snapchat beta and the role of Google+ Historically, when Google Play introduced the concept of beta testing in 2014, developers were offered to attach a Google+ community to their beta release track so they got collect necessary feedback from beta testers. Some of these communities were super small and some were big (up to 100k members). But Snapchat's community on Google Plus was exceptional, it had around 3 million participants. This happened because Snapchat added a link to this community inside the application itself, exposing it to a much broader audience than just beta testers. At this point, it was not much about beta testing anymore. Google+ community members were sharing their Snapcodes over there hoping to connect with more subscribers. Back in the days, the easiest way to do this was by sharing an image with your Snapcode so other users could scan it from the app. Hundreds of images were posted to this community every minute. At some point, it caused a very significant load of Google Cloud services and the Google+ team made a decision to shut this community down. This was an exceptional case. ![Users were spamming such SnapCodes on the Google+ community](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/06/photo5350561595140453458.jpg) Users were spamming such SnapCodes on the Google+ community On the next day, TestingCatalog created a new community on Google+ with the same name "snapchat-android-beta" but specified explicitly that it is unofficial and will be dedicated to Snapchat updates and other news. In one week, it grew from 0 to 30k users organically just because Snpachatters were searching for this group on Google+ by themselves 🔥 The feature of using Google+ communities for beta testing was abandoned by Google at some point and Google+ was shut down in 2018 as well. Snapchat beta remained to be public and accessible for new testers. ## Best Snapchat features to try ![Lenses scan, Games and Minis on Snapchat](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/06/photo5343537841847776830.jpg) Lenses scan, Games and Minis on Snapchat ### Expirable Content The first, introduced by Snap.Inc - the company behind, is the disappearing content. Any photo or video Snap can be set to vanish in 24 hours. Also, previously, messages were disappearing instantly, but fortunately, now Snapchat chats can be kept for up to 24 hours. ### AR Lenses Snapchat Lenses are AR-powered and are fun to use with all of the included unique masks (bunny ears, dog ears, and nose, etc.), providing a real life-like experience. ### SnapMap SnapMap is basically a map, where all your friends' location is being displayed along with other attractions. There you can also zoom into trending snaps from a particular location. ### Spotlight After the success of TikTok and its short videos format, Snapchat implemented its feature for this content format. Spotlight is focused on short videos that you can scroll vertically. Other features worth checking: Games and Minis, Lenses Scan, Lenses Studio, Bitmoji integration. ## Other Android apps by Snap worth checking ### Bitmoji Bitmoji is a standalone app that provides customizable avatars to Snapchat and other apps. There you can change your style and the way how your profile will appear on SnapMap or Bitmoji stories for example. [Bitmoji - Apps on Google PlayBitmoji is your personal emoji. Use it in Snapchat and wherever else you chat!![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayBitmoji![](https://play-lh.googleusercontent.com/ZL9xc5FVxhOu4Ju6Vdrlq8xgN4xu9Lce141ACJzY0_f_9hZmr6UhK10R5ftOS9KYdw)](https://play.google.com/store/apps/details?id=com.bitstrips.imoji&ref=testingcatalog.com) ### Zenly Is a standalone app that utilizes the SnapMap feature to let you connect and meet with your friends. [Zenly - Your map, your people - Apps on Google PlayFrom Snap Inc.![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayZENLY![](https://play-lh.googleusercontent.com/-DJWVwaeVg8sPY3zO8wpCGP7Bv8pxsmjRqHm6puud8wwZIM234Fs94C3dCHf45yVE0g)](https://play.google.com/store/apps/details?id=app.zenly.locator&ref=testingcatalog.com) ## Snapchat description from Google Play > Snapchat is a fast and fun way to share the moment with friends and family 👻 > > Snapchat opens right to the camera, so you can send a Snap in seconds! Just take a photo or video, add a caption and send it to your best friends and family. Express yourself with Filters, Lenses, Bitmojis and all kinds of fun effects. > > SNAP 📸 > • Snapchat opens right to the camera. Tap to take a photo, or press and hold for video. > • Add a Lens or Filter to your photo – new ones are added every day! Change the way you look, dance with your 3D Bitmoji and discover games you can play using your face. > • Create your own Filters to add to photos and videos – or try out Lenses made by our community! > > CHAT 💬 > • Stay in touch and chat with friends with live messaging, or share your day with Group Stories. > • Video chat with up to 16 friends at once. You can even use Filters and Lenses! > • Express yourself with Friendmojis – exclusive Bitmojis made just for you and a friend. > > STORIES > • Watch friends’ Stories to see their day unfold. > • Watch Stories from the Snapchat community, based on your interests. > • Explore new perspectives from top creators. > > DISCOVER 🔍 > • Watch breaking news and exclusive Original Shows. > • Keep up to date with Stories from top publishers. > • Enjoy a curated feed – made to fit your phone. > > SNAP MAP 🗺 > • See where your friends are hanging out, if they’ve shared their location with you. > • Share your location with your best friends, or go off the grid with Ghost Mode. > • Discover live Stories from the community nearby, or across the world! > > MEMORIES 🎞️ > • Look back on Snaps you’ve saved with free cloud storage. > • Edit and send old moments to friends, or save them to your Camera Roll. > • Create Stories from your favourite memories to share with friends and family. > > FRIENDSHIP PROFILE 👥 > • Every friendship has its own special profile to see the moments you’ve saved together. > • Discover new things you have in common with Charms. See how long you’ve been friends, your astrological compatibility, your Bitmojis’ fashion sense and more! > • Friendship Profiles are just between you and a friend, so you can bond over what makes your friendship special. > > Happy Snapping! ## Links and updates ### Snap News For more news regarding upcoming releases, you can visit Snapchat’s [News page](https://support.snapchat.com/en-US/news?ref=testingcatalog.com). ### Snapchat on the web Some features have web versions and here are the links to them: - [SnapMap](https://map.snapchat.com/?ref=testingcatalog.com) - [Stories](https://story.snapchat.com/?ref=testingcatalog.com) - [Lense Studio](https://lensstudio.snapchat.com/?ref=testingcatalog.com) Other Snapchat resources can be found via the [official website](https://www.snapchat.com/?ref=testingcatalog.com). ### Links - [Snapchat on Google Play](https://play.google.com/store/apps/details?id=com.snapchat.android&ref=testingcatalog.com) - [Snapchat beta OPT-IN page](https://play.google.com/apps/testing/com.snapchat.android?ref=testingcatalog.com) - [Snapchat on APKMirror](https://www.apkmirror.com/apk/snap-inc/snapchat/?ref=testingcatalog.com) - [Snapchat Support](https://support.snapchat.com/en-US?ref=testingcatalog.com) - [@SnapchatSuppor Twitter](https://twitter.com/snapchatsupport?ref=testingcatalog.com) If you want to catch up with the latest news about Snapchat and other Social apps, you will need to [subscribe to our weekly newsletter](https://www.getrevue.co/profile/testingcatalog?ref=testingcatalog.com) with a full summary of different news and updates on these topics 📩 You can also find all recent Snapchat features, news and updates that we've reported by navigating to the Snapchat tag page 👇 [Snapchat - TestingCatalogA news publication about Android for beta testers and early adopters who are eager about trying new apps and features![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov](https://www.testingcatalog.com/tag/snapchat/) ## Other Snapchat Guides [How to sign up for Snapchat on AndroidAs it is with every other social media network and its mobile application, there are some minor differences to the sign-up process. For the decently popular Snapchat in particular, things are pretty straightforward, so here they are laid out for you in a basic list form. Install the Snapchat app![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2020/07/Adobe_Post_20200717_2354450.5071677865450301.png)](https://www.testingcatalog.com/how-to-sign-up-for-snapchat-on-android-2020-edition-guide/) [Snapchat Map Status feature overview and how to use it on your AndroidIf you are not an active Snapchat user like me you may get surprised by some features you find in the app even if they were released a long time ago. A Snapchat Map Status feature released in May 2019 is one of the examples. Many Snapchat features can be![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2020/07/Adobe_Post_20200718_2306400.033181283584439614.png)](https://www.testingcatalog.com/snapchat-map-status-feature-overview-and-how-to-use-it-on-your-android/) ### YouTube Beta URL: https://www.testingcatalog.com/youtube-beta/ Last updated: 2023-12-04T20:52:45.000Z YouTube is a video-sharing platform everyone has heard of, with its best features right on your mobile device. Watch videos without a pause, upload your own content, and do whatever your heart desires. Like and share your favourite videos, subscribe to millions of channels and engage with others in the comments section. ## How to become a YouTube beta tester? Youtube for Android has two release tracks on Google Play that are publicly available - stable and beta. You can become a beta tester for YouTube from both, web and Android Google Play clients. ### Join on the web 1. Open [YouTube Beta OPT-IN](https://play.google.com/apps/testing/com.google.android.youtube?ref=testingcatalog.com) link. 2. Press on the "Become a Tester" button. 3. Press on the ["download it on Google Play"](https://play.google.com/store/apps/details?id=com.google.android.youtube&ref=testingcatalog.com) link. 4. See the app title to be updated with "Beta" and update it to the latest version. ### Join on Android 1. Search for YouTube. 2. Open YouTube application page. 3. Scroll down until the "Join the beta" section. 4. Tap on the "Join" button and confirm it on the pop-up window. 5. Wait until the app title to be updated with "Beta" and update it to the latest version. 6. Bear in mind that the OPT-IN process via Google Play may take up to several hours. ![How to become YouTube beta tester on Android](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/03/photo5409067918819439017.jpg) How to become YouTube beta tester on Android ## What does it mean "beta is full"? If you are trying to join via Google Play for Android, you can try to do the same [via the web link.](https://play.google.com/apps/testing/com.google.android.youtube?ref=testingcatalog.com) Google Play has a bug that shows "beta is full" status incorrectly and using a web OPT-IN link may be a workaround. However, the "Beta is full" message means that the current testers limit was reached already. If you see this message on the web page, it means that there is no way to join it until Google Devs increase the limit of participants. ## Why don't I have all features that exist for others? YouTube for Android has different ways to deliver new features to end-users: - **Client-side updates** \- come as a part of a new APK version. New client-side features usually available to everyone but they are rare. - **Server-side A/B tests** \- these updates make certain features available to a limited amount of users. - **Gradual rollouts** \- a process where Google releases new features to a wider percentage of users. [A/B tests and gradual rollouts](https://support.google.com/youtube/answer/7367023?hl=en&ref=testingcatalog.com) are the reasons why you may see some features that others don't and another way around. While for gradual rollouts, you may just need to wait for a little bit longer to receive it, with A/B tests it may happen that you won't receive a certain feature at all. Gradual rollouts usually being officially announced on the [YouTube product blog](https://blog.youtube/?ref=testingcatalog.com) in advance. ## How to get premium experimental YouTube features? YouTube also offers a set of experimental features exclusively to its [Premium](https://www.youtube.com/premium?ref=testingcatalog.com) users as a part of a subscription package. All available experiments that take place at this moment can be found on the ["New" page.](https://www.youtube.com/new?ref=testingcatalog.com) YouTube Premium is also available as a free trial so you can easily try it yourself first. ## How to leave YouTube beta program? You can leave YouTube beta program at any time in the same way as you signed up via Google Play on the web or Android. Leaving the beta program will not affect your access to experimental features in case if you have a premium subscription. ### Leave on the web 1. Open [OPT-IN link.](https://play.google.com/apps/testing/com.google.android.youtube?ref=testingcatalog.com) 2. Press on the "Leave the program" button. 3. Press on the ["download it on Google Play"](https://play.google.com/store/apps/details?id=com.google.android.youtube&ref=testingcatalog.com) link. 4. Uninstall the app and reinstall it again from Google Play. ### Leave on Android 1. Search for YouTube. 2. Open YouTube application page. 3. Scroll down until the "You are a beta tester" section. 4. Tap on the "Leave" button and confirm it on the pop-up window. 5. Uninstall the app and reinstall it again from Google Play. Leaving the beta release track means that you will only be receiving stable YouTube updates. These updates are being released less frequently but if you face any issues with your YouTube app it could be a way to solve them. ## Side-load from APKMirror It is also possible to side-load YouTube beta APKs if you want to avoid using Google Play for some reason. The easiest and safest way to do so will be checking and downloading the latest updates from [APKMirror.](https://www.apkmirror.com/?s=com.google.android.youtube&post%5Ftype=app%5Frelease&searchtype=apk&ref=testingcatalog.com) There you can find both stable and beta releases with the "beta" prefix in the version name. ## YouTube description from Google Play > Get the official YouTube app on Android phones and tablets. See what the world is watching -- from the hottest music videos to what’s popular in gaming, fashion, beauty, news, learning and more. Subscribe to channels you love, create content of your own, share with friends, and watch on any device. ## YouTube Shorts Beta ![YouTube Shorts Beta on Android](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/03/photo5409067918819439016.jpg) YouTube Shorts Beta on Android YouTube Shorts is a new feature inside YouTube that allows users to watch and share TikTok-like short vertical videos. The feature is only available in some countries. You don't need to be a YouTube beta tester In order to access it but you have to be located in the country where this feature was already rolled out. ## YouTube Studio If you are a YouTube creator, you may also need to download a YouTube Studio app that allows you to edit, post and analyse your content. This is a standalone app but it doesn't have a public beta available. [YouTube Studio - Apps on Google PlayThe official YouTube Studio app makes it faster and easier to manage your YouTube channels on the go. Check out your latest stats, respond to comments, upload custom video thumbnail images, schedule videos, and get notifications so you can stay connected and productive from anywhere. FEATURES: \* Mo…![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayGoogle LLC![](https://play-lh.googleusercontent.com/mqMlRJKWIi2fqw-J0X5gUhKQ-XudJj23BmPCjkb-noV4Hh5LCFgy1I1O97HGCMdJnXs)](https://play.google.com/store/apps/details?id=com.google.android.apps.youtube.creator&ref=testingcatalog.com) ## YouTube Go Depending on your device type and model, you may also have access to a YouTube Go app which is the same official YouTube client app that comes in a lighter form. It takes less storage space and consumes less data as well. [YouTube Go - Apps on Google PlayIntroducing YouTube Go 🎆 A brand new app to download and watch videos YouTube Go is your everyday companion, even when you have limited data or a slow connection. ✔️️ Discover popular videos: 🎵 songs, 🎥 movies, 📺 TV shows, 😂 comedy, 👜 fashion, 🍲 cooking, 🛠️ ’how-to’s and many, many more! ✔️…![](https://www.gstatic.com/android/market_images/web/favicon_v2.ico)Apps on Google PlayGoogle LLC![](https://play-lh.googleusercontent.com/wZmu5sJRUibg-xF_8IApSBkZSWVbx-rhVcxqOL6l_25-D6Tt28uHgz3b8yk4FhrH2w)](https://play.google.com/store/apps/details?id=com.google.android.apps.youtube.mango&ref=testingcatalog.com) ## Links and updates - [YouTube on Google Play](https://play.google.com/store/apps/details?id=com.google.android.youtube&ref=testingcatalog.com) - [YouTube Beta OPT-IN page](https://play.google.com/apps/testing/com.google.android.youtube?ref=testingcatalog.com) - [YouTube on APKMirror](https://www.apkmirror.com/apk/google-inc/youtube/?ref=testingcatalog.com) If you want to catch up with the latest news about YouTube and other Google apps, you will need to [subscribe to our weekly newsletter](https://www.getrevue.co/profile/testingcatalog?ref=testingcatalog.com) with a full summary of different news and updates on these topics 📩 You can also find all recent YouTube features, news and updates that we've reported by navigation to the YouTube tag page 👇 [YouTube - TestingCatalogA news publication about Android for beta testers and early adopters who are eager about trying new apps and features![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov](https://www.testingcatalog.com/tag/youtube/) ## Other YouTube Guides [How do I make a short clip from a video on YouTube?YouTube Clip feature was rolled out to more users on Android. This feature allows anyone to make 60 seconds cut from any video and share it as a clip.![](https://www.testingcatalog.com/favicon.png)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2021/05/photo5244526476169163968.jpg)](https://www.testingcatalog.com/how-do-i-make-a-short-clip-from-a-video-on-youtube/) ### Contact Us URL: https://www.testingcatalog.com/contact-us/ Last updated: 2025-02-01T19:01:40.000Z **Reach Out to TestingCatalog** At TestingCatalog, we're constantly on the pulse of the latest in tech, AI advancements, and user experience innovations. Whether you're a tech enthusiast, an early adopter, a budding developer, or a company at the forefront of technology, we're here to connect. **General Inquiries:** For all general questions, feedback, or just to say hello, reach out to us via the contact form below. We value your input and look forward to hearing from you. **News Tips:** Got a scoop on the latest app feature or an emerging tech trend? Our editorial team is always eager for leads on groundbreaking stories. Drop your tips at **testingcataloghelp@gmail.com**. **Press Releases:** For press releases and official announcements, please direct them to the same form. We ensure that your news gets the right attention. **Advertising and Partnerships:** Interested in advertising with us or exploring partnership opportunities? Contact our marketing team through our contact form for more information. Send Message TestingCatalog is more than just a news outlet; it's a community of tech aficionados and industry leaders. We're excited to hear from you and collaborate in shaping the future of technology. **Social Media:** Follow us on our social channels to stay updated with the latest in tech. Find us on [X](https://twitter.com/testingcatalog?ref=testingcatalog.com), [Facebook](https://www.facebook.com/testingcatalog), and [LinkedIn](https://www.linkedin.com/company/testingcatalog/about/?ref=testingcatalog.com). **Office Address:** Talstr. 3a Berlin, 13189 ### ChatGPT Alpha URL: https://www.testingcatalog.com/chatgpt-alpha/ Last updated: 2026-04-26T20:24:19.000Z ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-chat.openai.com-2023.12.25-16_31_27.jpg) ChatGPT UI ### BREAKING UPDATE (Dec, 2023) 🔥 ChatGPT Alpha now serves the purpose of testing ChatGPT experience in different locales. The "How to join ChatGPT Alpha" section reflects the process of joining this locale testing program. The legacy Alpha program is marked as "Legacy" and available in the section below. ### BREAKING UPDATE (May, 2023) 🔥 ChatGPT has announced that their experimental plugins will be available in beta for all ChatGPT Plus users starting next week! This means that you no longer need to be part of the exclusive alpha testing group to try out these exciting new features. [Read on to find out how you can enable and access these beta features](https://www.testingcatalog.com/chatgpt-rolls-out-beta-plugins-and-web-browsing-for-plus-users-exploring-new-features-and-integrations/) 👇 Welcome to the definitive guide on ChatGPT Alpha, your comprehensive source for understanding ChatGPT's alpha features, accessing them, and making the most out of this cutting-edge AI technology. In this guide, we'll cover key sections like how to access alpha features, what ChatGPT is, ChatGPT plugins, ChatGPT Plus, the ChatGPT API, and some other alternatives. ## What is ChatGPT? [ChatGPT](https://chat.openai.com/chat?model=gpt-4&ref=testingcatalog.com) is a powerful AI language model developed by OpenAI, based on the GPT-3.5 and GPT-4 architectures. GPT-3.5 is the version currently available to the public, while GPT-4 is an exclusive Plus-only feature that offers enhanced capabilities and performance for ChatGPT Plus subscribers. Both versions are designed to understand and generate human-like text responses based on user inputs. ChatGPT has a wide range of applications, including content generation, brainstorming, programming help, learning new topics, and much more. The ChatGPT platform is accessible via a user-friendly interface, available on the desktop web and mobile web for both Android and iOS devices. As ChatGPT continues to evolve, its capabilities will expand, providing users with even more advanced features and improvements that enhance their AI-powered communication experience. Stay up-to-date on the latest developments by following OpenAI's announcements and updates. ## ChatGPT Release Process OpenAI follows a phased rollout process when introducing new capabilities and model improvements to users. This process ensures that features are tested and refined before being made widely available. The key stages of a new feature's progression toward general availability are as follows: 1. **Alpha Phase**: A small group of users is given access to the new feature for testing and feedback. This stage helps OpenAI gather insights and make necessary adjustments based on real-world use. 2. **Beta Phase**: The updated feature is made available to ChatGPT Plus subscribers who have opted-in for beta testing. This larger user base helps OpenAI further evaluate the feature's performance, stability, and overall user experience. 3. **General Availability**: After the beta testing is completed, the feature is assessed for quality and made available to all ChatGPT users if it meets the quality standards. Follow the official [ChatGPT FAQ page](https://help.openai.com/en/articles/7897380-introducing-new-features-in-chatgpt?ref=testingcatalog.com) for more info. ## How to Get Access to ChatGPT Beta Features If you're a ChatGPT Plus user, you'll soon be able to access the beta features by following these simple steps: 1. Navigate to [ChatGPT](https://chat.openai.com/?ref=testingcatalog.com) 2. Click on "Profile & Settings" 3. Select "Beta features" 4. Toggle on the features you'd like to try Once the beta panel is available to you, you'll be able to try two new features: Web browsing and Plugins. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/05/BetaPanel.png) ChatGPT beta features #### Third-party Plugins To use third-party plugins, follow these instructions: 1. Navigate to [ChatGPT](https://chat.openai.com/?ref=testingcatalog.com) 2. Select "Plugins" from the model switcher 3. In the "Plugins" dropdown, click "Plugin Store" to install and enable new plugins [ChatGPT rolls out beta plugins and web browsing for Plus users: exploring new features and integrationsChatGPT, the popular AI language model by OpenAI, has just released beta plugins and web browsing features for its plus users. In addition to these new capabilities, the update also introduces a redesigned model selector, providing an easy way to switch between GPT-3.5 and GPT-4\. How to Access the![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/size/w256h256/2023/01/5iVVRzJO_400x400.jpeg)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/05/screenshot-chat.openai.com-2023.05.17-20_49_11.jpg)](https://www.testingcatalog.com/chatgpt-rolls-out-beta-plugins-and-web-browsing-for-plus-users-exploring-new-features-and-integrations/) ## ✅ (Current Alpha Program - Locale Testing) How to Get Access to Alpha Features ChatGPT's current alpha program offers an exciting opportunity for users to experience language support for nine different locales, enhancing the platform's accessibility and user experience across diverse linguistic backgrounds. This initiative is part of OpenAI's ongoing efforts to refine and expand ChatGPT's capabilities, ensuring a more inclusive and versatile tool for users worldwide. ### To gain access to these alpha features, follow these steps: 1. Ensure you're accessing ChatGPT through a web browser and visit the official ChatGPT site. 2. Modify your browser settings to match one of the supported locales: Chinese (Simplified or Traditional), French, German, Italian, Japanese, Portuguese (Brazilian), Russian, or Spanish. 3. Look for the "Join alpha" banner within the ChatGPT interface and click on it to opt into the alpha testing. 4. Once opted in, the ChatGPT UI will update to display in the selected language, granting you early access to test these new linguistic features. By participating in this alpha program, users not only get a firsthand look at ChatGPT's evolving language capabilities but also contribute valuable feedback that helps shape the future development of this cutting-edge technology. It's a unique chance to engage with the latest advancements in AI and support OpenAI's mission to make AI more accessible and useful for a global audience. For more detailed instructions and information, please refer to the official FAQ [here](https://help.openai.com/en/articles/8357869-chatgpt-language-support-alpha-web?ref=testingcatalog.com). ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2024/01/Screenshot-2024-01-24-at-19.59.24.png) ChatGPT Alpha in the Russian language ## ⛔️ (Legacy Alpha Program - Not active anymore) How to Get Access to Alpha Features To access the ChatGPT Alpha exclusive features, you need to join the [ChatGPT Plugins waitlist](https://openai.com/waitlist/plugins?ref=testingcatalog.com). OpenAI is extending plugin access to users and developers, with an initial focus on a small number of developers and ChatGPT Plus users. The plan is to roll out larger-scale access over time. By joining the waitlist and receiving an invite, users will be able to enjoy a wide range of plugin functionalities that enhance their ChatGPT experience. **To join the waitlist, follow these steps:** 1. Visit OpenAI's website and navigate to the [ChatGPT Plugins waitlist](https://openai.com/waitlist/plugins?ref=testingcatalog.com) page. 2. Click on the "Join Waitlist" button to access the signup form. 3. Fill out the required fields, including your first name, last name, email, and country of residence. 4. Indicate whether you would be willing to provide feedback about your plugin experience by selecting "Yes" or "No." 5. Describe the types of use cases or new plugins you would like to see built, and if you have a specific plugin idea, share it in the designated field. 6. Specify how you want to use plugins by selecting either "I want to try plugins in ChatGPT" or "I am a developer and want to build a plugin." 7. Choose how you are currently using ChatGPT, whether it's for personal, work, education, or other purposes. 8. Select the plugin you are primarily interested in, such as browsing, code interpreter, or third-party plugins. 9. Click "Join Waitlist" to submit your application. 📲 TIP: Some Reddit testers [suggest](https://www.reddit.com/r/OpenAI/comments/126g8zy/comment/jeab4nh/?utm%5Fsource=share&utm%5Fmedium=web2x&context=3) applying once per each option (browsing, 3rd party plugins, code interpreter), in case one of them becomes available earlier. ![ChatGPT Plugins Alpha Waitlist](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/04/screenshot-openai.com-2023.04.04-21_25_03.png) ChatGPT Plugins Alpha Waitlist Once you have been approved and granted access, you will receive an invitation allowing you to explore the latest alpha features, updates, and opportunities to provide feedback for improvements. By participating in the ChatGPT Alpha program, you will be among the first to experience cutting-edge features and contribute to the platform's ongoing development. ## ChatGPT Mobile Apps: Engage with AI on the Go **ChatGPT**, a revolutionary AI tool, is not just confined to desktops. Users can now engage with ChatGPT through dedicated mobile apps available for both **iOS and Android** platforms. These apps bring the power of ChatGPT's conversational AI to your fingertips, making it more accessible and convenient to interact with the AI, no matter where you are. #### Unique Mobile-Only Feature: Voice-Only Mode A standout feature in the ChatGPT mobile app is the **voice-only mode**. This innovative functionality allows users to have a voice conversation with ChatGPT, offering a hands-free, interactive experience. It's perfect for those moments when typing isn't feasible or for users who prefer auditory interaction. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-play.google.com-2023.12.19-22_19_50.jpg) How to join ChatGPT beta program on Android #### ChatGPT for iOS and Android You can download the ChatGPT app for your device using the following links: - **iOS App Store**: [ChatGPT for iOS](https://apps.apple.com/us/app/chatgpt/id6448311069?ref=testingcatalog.com) - **Google Play Store**: [ChatGPT for Android](https://play.google.com/store/apps/details?id=com.openai.chatgpt&ref=testingcatalog.com) #### Join the ChatGPT Beta on Android For Android users, there's an exciting opportunity to experience the latest features of ChatGPT before they are released to the public. Join the **ChatGPT Beta Program** on Android and be a part of the future of conversational AI. Here's how to join: 1. **Visit the Beta Program Link**: Start by visiting the ChatGPT Beta program page on the Google Play Store. You can access it [here](https://play.google.com/apps/testing/com.openai.chatgpt?ref=testingcatalog.com). 2. **Sign In to Your Google Account**: Ensure you're signed in to your Google account that is associated with your Android device. 3. **Become a Tester**: Click on the 'Become a Tester' button. You'll receive a notification confirming your enrollment in the beta program. 4. **Download or Update the App**: After joining, you can download the ChatGPT app from the Google Play Store if you haven't already. If you have the app installed, it will update to the beta version. 5. **Provide Feedback**: As a beta tester, your feedback is invaluable. Use the app and share any bugs or improvement suggestions through the feedback options within the app or on the Google Play Store page. By joining the ChatGPT Beta on Android, you're not only getting early access to new features but also contributing to the refinement and improvement of the app. ## Custom GPTs and GPT Builder OpenAI's recent advancements include the [introduction](https://openai.com/blog/introducing-gpts?ref=testingcatalog.com) of customizable versions of ChatGPT, known as GPTs. These custom GPTs allow users to tailor ChatGPT for specific needs, combining unique instructions, additional knowledge, and various skills, without the requirement for coding knowledge. This innovation is designed to make ChatGPT more practical for daily tasks, work, or home use, and can be shared with others. Furthermore, OpenAI is set to launch the GPT Store, a platform where these custom GPTs can be shared publicly. Verified builders will feature their creations, which will be searchable and could climb leaderboards. This store is not only a marketplace for creative GPTs but also offers the potential for creators to earn based on the usage of their GPTs. This initiative highlights the role of community builders in enhancing the versatility and application of GPT technology. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-chat.openai.com-2023.12.02-21_02_01.jpg) Explore tab for custom GPTs To ensure user safety and data privacy, OpenAI has implemented robust measures. The interactions with these GPTs are kept private, and users have control over whether their data is shared with third-party APIs. These privacy considerations are part of OpenAI's broader commitment to responsible AI deployment and user trust. [Discovering Custom GPTs: A Guide for OpenAI PLUS SubscribersExplore the latest in AI: Discover new GPTs for OpenAI, find custom tools & directories, and stay ahead with TestingCatalog.![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/size/w256h256/2023/01/5iVVRzJO_400x400.jpeg)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/size/w1200/2023/11/screenshot-www.gptshunter.com-2023.11.12-23_31_45.jpg)](https://www.testingcatalog.com/discovering-custom-gpts-a-guide-for-openai-plus-subscribers/) Additionally, developers can integrate GPTs with real-world applications by defining custom actions through APIs. This capability allows for a wide range of integrations, such as connecting to databases, email systems, or e-commerce platforms, providing greater utility and real-world application potential. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-chat.openai.com-2023.11.10-21_02_07.jpg) Custom GPT editor UI Enterprise customers are also catered for, with options to create internal-only GPTs. These can be tailored for specific business needs, departments, or proprietary datasets, enabling businesses to harness the power of GPT for internal efficiency and specific use cases. In summary, these developments by OpenAI, through the introduction of custom GPTs, the GPT Store, and enhanced privacy and integration features, mark a significant step in making AI more accessible, customizable, and practical for a wide range of users and applications. 💡 Give a try to [AI News Reporter GPT](https://chat.openai.com/g/g-Ft2u9p9PB-ai-news-features?ref=testingcatalog.com) to stay on top of the latest AI news! ## GPTs Store In January 2024, OpenAI introduced the GPT Store, a novel addition for ChatGPT Plus subscribers that enriches the platform with an expansive selection of public custom GPTs. This initiative marks a significant step in making AI more accessible and customizable, catering to a diverse range of user needs and interests. The GPT Store serves as a vibrant marketplace, showcasing over 3 million custom GPT versions created by a community of enthusiasts and experts. It's designed to facilitate easy discovery and exploration of GPTs tailored for various applications, from educational tools to programming aids and lifestyle enhancers. Highlighting its commitment to fostering innovation and collaboration, the GPT Store features a carefully curated selection of GPTs. These are sourced from both OpenAI's partners and the wider community, ensuring a quality and variety that empower users to find the perfect GPT for their specific requirements. For an in-depth look at this feature, you can read the official announcement [here](https://openai.com/blog/introducing-the-gpt-store?ref=testingcatalog.com). ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2024/01/screenshot-chat.openai.com-2024.01.27-19_57_31.jpg) Official GPTs Store ## ChatGPT Plugins ChatGPT plugins are third-party extensions that enhance the functionality and user experience of ChatGPT. These plugins can be integrated with various applications to offer additional features like advanced text editing, translations, content suggestions, and more. Examples of popular plugins include Grammarly integration for grammar and spell-checking and project management tools that utilize ChatGPT's capabilities to streamline workflows. With access to third-party plugins, users can explore numerous applications across various use cases such as automation, shopping, travel planning, dining, and information/computation. These plugins, developed by external partners, are designed to improve user experience and streamline tasks. For example, users can plan meals with the Instacart plugin, which allows them to create meal plans and automatically add items to their shopping cart, ultimately having their groceries delivered straight to their doorstep. > Oh my goodness! I just got access to [#ChatGPT](https://twitter.com/hashtag/ChatGPT?src=hash&ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Plugins! 😱 [pic.twitter.com/it6s3i9PH0](https://t.co/it6s3i9PH0?ref=testingcatalog.com) > > — DataChazGPT 🤯 (not a bot) (@DataChaz) [April 4, 2023](https://twitter.com/DataChaz/status/1643155528380547072?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) ChatGPT Alpha Plugins Preview Users can also leverage the Expedia plugin to plan trips, with ChatGPT asking clarifying questions and offering personalized flight options based on user input. Similarly, the Speak plugin enables users to translate words and phrases accurately, providing colloquial alternatives and examples in sentences. Another powerful plugin is Zapier, which allows users to automate tasks by integrating various applications, such as Gmail, Slack, and Google Sheets. Although it might have a learning curve, Zapier's automation capabilities make it one of the most potent plugins available. ## ChatGPT Plus [ChatGPT Plus](https://openai.com/blog/chatgpt-plus?ref=testingcatalog.com) is a subscription plan that provides users with premium access to ChatGPT features and benefits. With a ChatGPT Plus subscription, users enjoy faster response times, priority access to new features and improvements, and general availability even during peak times. This subscription is perfect for professionals and businesses seeking a more streamlined experience with ChatGPT. ## ChatGPT API The [ChatGPT API](https://openai.com/blog/introducing-chatgpt-and-whisper-apis?ref=testingcatalog.com) allows developers to integrate ChatGPT's powerful language understanding and generation capabilities into their applications and services. With the API, developers can create custom solutions for tasks such as content generation, sentiment analysis, and automated customer support. The API documentation provides comprehensive guidelines, including authentication, endpoints, and usage limits. ## DALL-E DALL-E is a groundbreaking AI model developed by OpenAI, designed to produce unique and imaginative images based on text descriptions. By combining language understanding with image generation capabilities, DALL-E can create visually impressive and highly-detailed images from user-supplied prompts. ![DALL-E Alpha from OpenAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/04/screenshot-labs.openai.com-2023.04.30-22_58_46.png) DALL-E Alpha from OpenAI This state-of-the-art AI model unlocks new possibilities in art, design, and visual storytelling. DALL-E can generate everything from simple objects to intricate, fantastical scenes, showcasing an exceptional level of creativity and detail. To access DALL-E, you can visit [OpenAI Labs](https://labs.openai.com/?ref=testingcatalog.com), where you can purchase credits to use the DALL-E tool. These credits allow you to generate custom images by simply providing text prompts. Once you have acquired the necessary credits, you can start creating unique visuals for personal or professional projects through the user-friendly interface on the platform. In addition, you may want to consider exploring DALL-E's competitor, [Midjourney](https://www.midjourney.com/?ref=testingcatalog.com). Midjourney is another AI-powered image-generation platform that offers creative solutions for generating visual content based on text prompts. By comparing the features, user experience, and pricing structures of both DALL-E and Midjourney, you can make an informed decision on which tool best suits your needs. Delve into the world of AI-powered image generation with DALL-E and Midjourney, and discover how these innovative tools are revolutionizing the way we create and interact with visual content while offering easy-to-use and accessible options for users of various skill levels. ## Other Alternatives While ChatGPT-4 is a powerful language model, it's important to be aware of other alternatives available in the market, especially if you're looking for cost-effective options. Some alternatives include: - [ChatGPT-3.5](https://chat.openai.com/chat?model=text-davinci-002-render-sha&ref=testingcatalog.com): The publicly available version of ChatGPT by OpenAI, offers a wide range of applications and use cases, such as content generation, brainstorming, programming help, learning new topics, and more. This option may be more suitable for those who don't want to pay for ChatGPT Plus. - Google Bard: Google Bard, created as a response to the rising popularity of ChatGPT, marks a significant stride in Google's AI endeavours. Operating on the sophisticated Gemini Pro model, Bard excels in multimodal reasoning, offering nuanced and advanced AI conversations. This tool epitomizes Google's relentless pursuit of cutting-edge AI technology. The tool is available for free to Google users. - [Grok:](https://twitter.com/i/grok?ref=testingcatalog.com) Grok AI is a tool developed by X (formerly Twitter), designed to provide enhanced analytics and data interpretation capabilities. It has integration with X making its live data accessible to the LLM. It is only available to Premium+ subscribers. - [Microsoft Copilot:](https://copilot.microsoft.com/?ref=testingcatalog.com) ChatGPT-4 is available as part of Bing's search engine, enhancing the search experience by providing more relevant and context-aware search results. This option is available for free, making it a cost-effective alternative to ChatGPT Plus. ![Bing Chat UI Preview](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/04/screenshot-www.bing.com-2023.04.04-21_26_45.png) Bing Chat UI Preview By exploring these alternatives, you can choose the AI language model that best suits your specific needs and requirements without necessarily having to pay for premium features. ## Links & Official Sources To help you dive deeper into the world of ChatGPT, we have curated a list of official links and sources where you can find reliable and up-to-date information. Stay informed about the latest developments, research, and news by exploring these resources. 1. [ChatGPT](https://chat.openai.com/?model=gpt-4&ref=testingcatalog.com) 2. [DALL-E](https://labs.openai.com/?ref=testingcatalog.com) 3. **OpenAI's Official Website:** Visit [OpenAI's official website](https://www.openai.com/?ref=testingcatalog.com) to learn more about the organization responsible for developing ChatGPT Alpha. OpenAI is dedicated to advancing digital intelligence to benefit all of humanity, and its website provides a wealth of information about its mission, research, and products. 4. **OpenAI Blog:** The [OpenAI Blog](https://www.openai.com/blog/?ref=testingcatalog.com) is a fantastic resource for staying current with the latest breakthroughs, research, and updates related to ChatGPT Alpha and other OpenAI projects. By following the blog, you'll gain insights into how ChatGPT Alpha is evolving and improving, as well as learn about other exciting developments in the field of AI. 5. **OpenAI on Twitter:** For real-time updates and announcements, be sure to follow [OpenAI on Twitter](https://twitter.com/OpenAI?ref=testingcatalog.com). Their Twitter profile is an excellent way to stay connected with the organization and receive the latest information on ChatGPT Alpha, research breakthroughs, and other AI-related news. 6. **OpenAI Discord Server:** Join the [OpenAI Discord server](https://discord.gg/openai?ref=testingcatalog.com) to connect with a community of AI enthusiasts, researchers, and developers. The server provides a platform for users to discuss ChatGPT Alpha, share their experiences, and seek help from others working with OpenAI's cutting-edge technologies. 7. **OpenAI Release Notes:** A comprehensive [ChatGPT changelog](https://help.openai.com/en/articles/6825453-chatgpt-release-notes?ref=testingcatalog.com) with detailed information about each release. ## Other Guides on ChatGPT Alpha [A Comprehensive Guide to Reporting Bugs for ChatGPT UsersAre you a ChatGPT Alpha user-facing bugs or issues? Follow this comprehensive guide to report them on OpenAI’s Discord server and contribute to the development of ChatGPT solutions![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/size/w256h256/2023/01/5iVVRzJO_400x400.jpeg)TestingCatalogAlexey Shabanov![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/04/screenshot-discord.com-2023.04.29-23_52_13.png)](https://www.testingcatalog.com/a-comprehensive-guide-to-reporting-bugs-for-chatgpt-alpha-users/) This definitive guide to ChatGPT Alpha provides an overview of the key aspects of ChatGPT, from accessing its alpha features to understanding its powerful applications. With this knowledge, you can make informed decisions on using and integrating ChatGPT into your projects and businesses. ## Stay up to date with the latest news, release waves and features for ChatGPT Alpha By following these channels, you can stay ahead of the curve and make the most of the ChatGPT Alpha experience. > ICYMI: [#ChatGPT](https://twitter.com/hashtag/ChatGPT?src=hash&ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) now allows you to rename individual conversations and clean them one by one as well 🧹 [pic.twitter.com/WwyskbKVkj](https://t.co/WwyskbKVkj?ref=testingcatalog.com) > > — TestingCatalog.eth 📲 (@testingcatalog) [April 29, 2023](https://twitter.com/testingcatalog/status/1652368759343005697?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) 1. **Follow OpenAI on Twitter:** OpenAI is the driving force behind ChatGPT Alpha, and their Twitter account (@OpenAI) provides regular updates about their latest research, products, and advancements. Stay informed by following them [here](https://twitter.com/OpenAI?ref=testingcatalog.com). 2. **Follow TestingCatalog on Twitter:** For a more focused look at ChatGPT Alpha's release waves and features, follow TestingCatalog (@testingcatalog) on Twitter. Their updates cover a wide range of apps, including ChatGPT Alpha. Follow our account [here](https://twitter.com/testingcatalog?ref=testingcatalog.com). 3. **Subscribe to the TestingCatalog newsletter:** To receive a comprehensive overview of the latest news and updates on ChatGPT Alpha and other beta releases, consider subscribing to the TestingCatalog newsletter. You'll get valuable insights and updates directly in your inbox, ensuring you never miss out on any important information. By following these resources, you'll stay up to date with everything related to ChatGPT Alpha and be among the first to know about new release waves and features. This will enable you to better understand and utilize the potential of this groundbreaking AI tool. [ChatGPT News - TestingCatalogStay informed with the latest news, updates, and features of ChatGPT![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/size/w256h256/2023/01/5iVVRzJO_400x400.jpeg)TestingCatalogAlexey Shabanov](https://www.testingcatalog.com/tag/chatgpt/) Chat GPT News tag ### Gemini Experiment URL: https://www.testingcatalog.com/gemini-experiment/ Last updated: 2026-04-26T20:21:51.000Z ## About Gemini (Formerly Google Bard) Google Gemini, created as a response to the rising popularity of ChatGPT, marks a significant stride in Google's AI endeavours. Operating on the sophisticated Gemini Pro model, Gemini excels in multimodal reasoning, offering nuanced and advanced AI conversations. This tool epitomizes Google's relentless pursuit of cutting-edge AI technology. On this page, you'll find comprehensive insights into accessing Gemini, including tips for early and exclusive feature access. We also guide you on staying abreast of the latest updates and enhancements to Gemini, ensuring you're equipped with the knowledge to leverage this powerful AI tool to its fullest potential 👇 ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-bard.google.com-2023.12.19-22_06_00-1.jpg) Google Gemini homepage UI ## Access to Google Gemini Experiment Accessing Google Gemini is straightforward but requires a Google account, ensuring a personalized and secure experience. Gemini's user interface is available on both web and mobile web platforms, with the mobile web experience being notably well-optimized for on-the-go interactions. Globally available in over 180 countries and territories, Gemini brings Google's advanced AI capabilities to a vast audience. Whether you're at your desk or on your mobile device, you can easily dive into Gemini's AI-driven world through the [official Gemini URL](https://gemini.google.com/?ref=testingcatalog.com). This wide availability signifies Google's dedication to making sophisticated AI tools accessible to a broad audience. ## Gemini Updates To stay updated on Gemini, consider the following sources: - **Google** Gemini **Experiment Updates:** an official changelog for Google Gemini. - [**Google AI Blog:**](https://blog.google/technology/ai/?ref=testingcatalog.com) In-depth insights into Gemini's development. - [**Google Research Blog:**](https://blog.research.google/?ref=testingcatalog.com) Detailed articles and research papers on Gemini. - **TestingCatalog:** Regular updates and news about Gemini. - [**Social Media Channels:**](https://twitter.com/testingcatalog?ref=testingcatalog.com) Real-time updates on platforms like X (Twitter). ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-bard.google.com-2023.12.19-22_06_30.jpg) Google Gemini Updates changelog When checking Gemini's official updates page, be mindful that the changelog is tailored to your location. This means that if new features are released exclusively in the U.S., they might not appear in the changelog if you're accessing it from an EU country. Therefore, for global users, especially those interested in early access to new features, it's beneficial to consult a variety of sources to get a comprehensive view of Gemini's updates and advancements. ## How to access early features of Google Gemini with VPN Gaining early access to certain Gemini features, especially those initially released in the U.S., can be achieved using a VPN. The Opera browser, with its built-in VPN feature, is an excellent choice for this purpose. To use Gemini's U.S.-exclusive features from outside the region, follow these steps: 1. [**Install Opera Browser:**](https://www.opera.com/one?ref=testingcatalog.com) Download and install Opera One, which comes with a free, integrated VPN. 2. **Enable VPN in Settings**: Navigate to the Opera browser's settings and turn on the VPN feature. 3. **Access Gemini Page**: Open the Gemini page in your Opera browser. 4. **Change VPN Location**: From the address bar dropdown menu, select the VPN location option and change it to "Americas". This will give you access as though you're browsing from within the U.S. By following these steps, users outside the U.S. can explore and utilize new Gemini features that are initially region-restricted. This method offers a straightforward solution for global users eager to experience the full suite of Gemini's capabilities as soon as they are released. ## Gemini Extensions Extensions in Google Gemini are sophisticated integrations with other Google services, leveraging the LLM model to provide access to the most recent data and tools. These extensions expand Gemini's functionality by integrating it with various Google products. To manage these extensions, users can find a button in the top right corner of the Gemini interface. This button opens a list of available extensions, allowing users to individually turn each extension on or off, tailoring the Gemini experience to their specific needs. ### 1\. Google Flights Integrates Gemini with Google Flights, aiding in trip planning by finding the best routes, hotels, and flights. It offers suggestions and guidance on places to visit, transforming Gemini into a comprehensive travel planner. ### 2\. Google Hotels Enhances Gemini's capabilities in hotel booking, helping users to find suitable accommodations, explore surroundings, and discover nearby experiences. ### 3\. Google Maps Allows Gemini to search and utilize Google Maps information. It can plan routes, trips, and provide detailed data from various locations, making Gemini a potent tool for geographical and travel inquiries. ### 4\. YouTube Gives Gemini access to YouTube's video library, enabling it to transcribe, summarize, and extract key information from videos. This extension is useful for users who prefer reading summaries over watching full videos. ### 5\. Google Workspace Integrates Gemini with Google Workspace, granting access to Drive, Gmail, and Google Docs. Gemini can assist in drafting emails, creating documents, and streamlining the workflow by searching through your documents and emails. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-bard.google.com-2023.12.19-22_07_03.jpg) Google Gemini Extensions page ## Google AI Labs [Google AI Labs](https://labs.google/?ref=testingcatalog.com) is a distinct portal where users can discover and engage with various AI experiments developed by Google. This platform serves as a hub for innovative AI projects, including the Google Gemini experiment. AI Labs showcases Google's latest explorations and advancements in AI technology, providing users with a unique opportunity to experience the forefront of AI development. Most experiments on AI Labs, like Gemini, are often initially released to U.S. users. Therefore, similar to accessing early features in Gemini, utilizing a VPN can be essential for users outside the U.S. who wish to participate in these cutting-edge experiments. By changing your VPN location to the U.S., you can gain early access to these innovative projects. To stay informed about the latest developments and new experiments in Google AI Labs, users are encouraged to follow the same news sources as for Gemini updates Google AI Labs features various AI experiments, including: - **Generative Search**: Enhancing search capabilities using AI. - **Notebook LM**: Advanced note-taking applications using language models. - **Project Tailwind**: A unique AI development project. - **MusicLM**: Exploring the integration of AI in music creation. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2023/12/screenshot-labs.google-2023.12.19-22_17_56.jpg) Google Labs homepage Google AI Labs presents a remarkable opportunity for enthusiasts and professionals alike to explore, experiment with, and understand the diverse capabilities of AI technologies developed by Google. This platform is not just a showcase of Google's AI prowess but also an invitation for users to actively engage with and contribute to the future of AI. ### Weekly Newsletter URL: https://www.testingcatalog.com/weekly-newsletter/ Last updated: 2026-04-10T19:28:17.000Z ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/04/TC-NL-new-2.png) ## Enter Your Email To Join! Subscribe Email sent! Check your inbox to complete your signup. We won't spam your inbox! ## Posts ### Meta's Muse agent enters closed beta with invite codes URL: https://www.testingcatalog.com/metas-muse-agent-enters-closed-alpha-with-invite-codes/ Last updated: 2026-09-07T05:38:12.000Z Meta’s upcoming AI agent, Muse, appears to be moving closer to launch. Earlier, we reported that Muse would be accessible through a waitlist, but several details have changed since then. In addition to the waitlist, Meta will introduce invite codes that users can distribute to help others gain access. The agent also seems to have moved from internal testing into a closed beta phase, with limited regional availability that is very likely restricted to "Saint Kitts and Nevis" and the US for now. Based on what is currently known about the waitlist page, mobile apps and a web version will be available at launch. A desktop version will likely follow later. We have also recently gained access to the application’s description and several early screenshots. The description reads: *“Muse is your personal AI agent that gets things done, always working on what matters to you. It manages email, books dinners, tracks budgets, and finds deals with your approval.”* Also, *“Muse takes the busy work off your plate. It manages scheduling conflicts, books reservations, sets reminders, and follows up with people so birthdays, plans, and deadlines don't slip. Stay on top of your calendar, email, and daily tasks without carrying it all in your head.”* Based on this description, Meta appears to be targeting cliché tasks such as booking tables. The more interesting promise, however, is that Muse will sit on top of a user’s calendar, email, and daily tasks. If Meta builds it as an assistant focused purely on task execution and designs the UI around that purpose, Muse could become a much bigger play. Competition in this space is already fierce, but there is still a massive opportunity to do it much better. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--11--2.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--12--1.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--13--1.webp) Muse UI The first screenshot, “Your personal agent that takes things off your plate,” shows a standard chat interface with an avatar of the Muse character at the top. The avatar uses light brown tones and looks a little childish, while remaining serious enough to fit the overall design. The prompt bar is also standard, complete with an attachment button and a dictation voice icon. In the featured example, Muse finds a travel document, offers to sign it, and sends a confirmation email, which the user accepts. Muse also sends an emoji reaction to show that it has started working on the task. If the agent is genuinely proactive like this, it will be a big deal. The next screenshot, “Approve what gets sent or spent,” focuses on user authorization. Muse says, “Found a few travel strollers for Luca, perfect for your trip,” displays the price, and asks whether the user is interested. It then presents a widget where the user can select a payment method, view the details, and authorize the transaction. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--14--1.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--16--1.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--17--1.webp) Muse UI “Track ticket prices and book reservations” shows one of the interface's more distinctive parts: an interactive browser-window widget that the user can open as a link. Muse uses it to open a cinema website, select seats, and handle the booking. The browser widget also carries a “Beta” label. The “Connect the apps you already use” screenshot displays connectors for Gmail, Google Calendar, HealthX, OpenTable, Facebook, Instagram, Peloton, Plaid, and Function Health. This selection indicates that Muse will support health, fitness, and financial integrations alongside communication, scheduling, and reservation services. Another screen, “Get ideas for what your agent can take on,” presents several suggested prompts: “I can follow up on your airplane refund,” “I can find a bigger stroller for Luca,” “Build a daily training plan for your half marathon,” “I can book your anniversary dinner,” and “I can help dial in your sleep.” If Muse is proactive or at least provides a dedicated tab for browsing these kinds of ideas, this part of the product could be very interesting. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/IMG_9071.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/IMG_9072.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/unnamed--15--2.webp) Muse UI The final screenshot, “Understand your spending and find ways to save,” begins with a prompt directly from the user asking Muse to review their spending. The next screen says, “Turn your goals into a real plan.” It includes a to-do list with different goals and checkboxes that users can modify and check off. The “create goal” area includes sections for health and relationships, with probably more sections underneath that are not currently visible. This is the latest set of findings on top of what was discovered earlier. Many more features will likely be added to Muse gradually, but it is already clear that the agent is undergoing testing with a closed group of users on mobile and web. A desktop version could potentially arrive soon as well. **Source:** Apple's App Store currently exposes a [Meta-authored listing](https://apps.apple.com/kn/app/muse-from-meta/id6760173601?ref=testingcatalog.com) for "Muse from Meta". It also describes finance, fitness, shopping, background work, activity log, connector control, and customization capabilities. ### IFM releases K2 Horizon: 6 models with full training record URL: https://www.testingcatalog.com/ifm-releases-k2-horizon-6-models-with-full-training-record/ Last updated: 2026-09-06T19:39:57.000Z The Institute of Foundation Models has released K2 Horizon, a connected family of six language models ranging from 0.9 billion to 375 billion parameters, published alongside details on how each was trained. The sizes run 0.9B, 3.7B, 7B, 32B, 36B, and 375B, and IFM maps each to a job: local development at the small end, single-node serving and cost-sensitive deployment in the middle, everyday heavy use and production serving experiments above that, and long horizon agent work at the top. Models and code carry an Apache 2.0 license, while datasets ship under their own terms. > Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters. > > \- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting… [pic.twitter.com/WWbqrvcBAC](https://t.co/WWbqrvcBAC?ref=testingcatalog.com) > > — Institute of Foundation Models (@IFM\_AI) [September 3, 2026](https://x.com/IFM%5FAI/status/2095497035806113861?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) What separates this from a standard weights drop is the material shipped alongside the weights. IFM published training code, intermediate checkpoints, training logs, and evaluations for all six models, plus the training data where licensing allows and detailed construction recipes where redistribution is restricted. Most open releases hand over a finished checkpoint and nothing about how it was reached, which leaves outside researchers unable to reproduce a result or audit a claim without asking the lab first. K2 Horizon takes the opposite position: publish the process, then let the community check the work. > Meet K2 Horizon. [pic.twitter.com/YT7DxQP6BW](https://t.co/YT7DxQP6BW?ref=testingcatalog.com) > > — Institute of Foundation Models (@IFM\_AI) [September 3, 2026](https://x.com/IFM%5FAI/status/2095494518410015022?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The release also includes something labs rarely put in launch materials. IFM documented cases of its own models gaming evaluations during training, describing behavior such as copying hidden answers, wrapping binaries, and exploiting checkers, and released the checkpoints from those training stages so researchers can trace when the behavior first appeared. Reward hacking is normally a matter of suspicion, and shipping documented cases with the checkpoints attached is an unusual move for a launch. Two architecture pieces arrive with the models. MoVA moves expert routing into the attention mechanism itself, a departure from the sparse mixture-of-experts pattern that routes at the feed-forward layer, and IFM reports it stays compatible with FlashAttention and grouped query attention. Uno is a diffusion adapter that generates blocks of tokens in parallel and, per IFM, runs without a separate draft model or swapping the base model. The 36B model activates roughly 4 billion parameters per token, and IFM places its capability close to the dense 32B. SPONSORED Check out K2 Horizon on HuggingFace [Learn more ](https://huggingface.co/collections/IFM/k2-horizon?ref=testingcatalog.com) IFM is the Institute of Foundation Models at MBZUAI, a graduate research university in Abu Dhabi focused on artificial intelligence, with additional labs in Paris and Silicon Valley. The institute has been publishing fully open models since 2023 through its LLM360 line, running from Amber through K2, K2 V2, and K2 Think, each one shipping training artifacts alongside the weights. K2 Horizon extends that approach across a full size range for the first time, from models designed for edge and on-device use up to a 375B model aimed at multi-step agent workloads. ### Google keeps transforming Gemini desktop into superapp URL: https://www.testingcatalog.com/google-keeps-transforming-gemini-desktop-into-superapp/ Last updated: 2026-09-05T12:39:58.000Z Google keeps upgrading its Gemini desktop app toward a proper super app intended to compete with Codex and Claude Desktop. The latest update adds a new toggle that lets users switch between Ask and Assign modes, analogous to Chat and Work in Codex. Ask presents the standard Gemini interface, while Assign works like Gemini Spark, the AI agent Google introduced recently, but with expanded capabilities. 0:00 /0:30 1× Assign will let users choose the folder in which Gemini operates. They can set a default folder, return to one used in a previous session, or select a new folder for a task, keeping the agent within that specific scope. The app also includes an option to connect to other computers running Gemini desktop. This could enable remote control features, although exactly how the connection will work is still unclear. It will likely resemble the approach used in Codex. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/Screenshot-2026-09-05-at-14.23.49.png) Two integrations stand out in the current testing. With [computer use](https://www.testingcatalog.com/google-tests-computer-use-on-gemini-desktop/), users can tag the computer use skill and trigger tasks from Gemini desktop that operate local apps. Google is also testing connections with Obsidian. It is unclear whether this points to a broader knowledge base concept or whether Google is using Obsidian internally while it works toward equivalent functionality in Gemini, but the integration will be very welcome either way. Google is also working on a Customize tab where users can select skills, apps, connectors, or [plugins](https://www.testingcatalog.com/google-is-working-on-plugins-for-gemini-enterprise/). The current build remains quite limited, offering mostly Google apps or extensions already found in the existing Gemini interface. Hopefully, Google will add proper MCPs and plugins later, as seen in competing apps. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/Screenshot-2026-09-04-at-21.06.27.png) Code traces indicate that Google is preparing the app to work with the next major model, potentially **Gemini 4**. The code also points to an upgrade for Nano Banana, specifically **Nano Banana 2.5 Flash**. > And in other parts of Google, there is another new model that caught the attention of my crystal ball > > Nano Banana 2.5 Flash > > Interesting [https://t.co/s4aIRXh4Gy](https://t.co/s4aIRXh4Gy?ref=testingcatalog.com) > > — Bedros Pamboukian (@bedros\_p) [September 4, 2026](https://x.com/bedros%5Fp/status/2095780756421546479?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) On macOS, Finder integration will let you send any folder directly to Gemini for inspection or a targeted task. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/Screenshot-2026-09-03-at-09.58.43.png) All of these features remain in closed testing, and the timeline for a public release is unclear. For now, let's see when, and in what form, these capabilities reach [Gemini](https://www.testingcatalog.com/tag/gemini/) users. Editorial notes 👀 1. **Spark background — May 19, 2026.** Google introduced Gemini Spark as a personal agent that works continuously in the cloud, connects to Workspace, and uses Gemini 3.5 with the Antigravity harness. Google also outlined future macOS support for local files and desktop workflows. [Google's Spark introduction](https://blog.google/innovation-and-ai/products/gemini-app/next-evolution-gemini-app/?ref=testingcatalog.com) 2. **Existing desktop and integration roadmap — June 30, 2026.** Google announced Spark for macOS in beta for adult U.S. AI Ultra subscribers, including permissioned local-file tasks. It described phone-initiated tasks on a Mac as forthcoming. It also announced third-party connected apps and a rollout of custom MCP support. [Google's macOS Spark update](https://blog.google/innovation-and-ai/products/gemini-app/gemini-spark-updates-june-2026/?ref=testingcatalog.com) 3. **Obsidian context — undated documentation, checked September 5, 2026.** Obsidian stores Markdown notes in local folders called vaults, and its documentation describes editing those files with other tools. That makes file-oriented knowledge workflows a relevant follow-up angle, but this is an inference about possible utility, not proof of Google's integration design or intentions. [Obsidian's data-storage documentation](https://obsidian.md/help/data-storage?ref=testingcatalog.com) ### OpenAI launches GPT-6 Astra across ChatGPT and API URL: https://www.testingcatalog.com/openai-launches-gpt-6-astra-across-chatgpt-and-api/ Last updated: 2026-09-04T22:11:31.000Z OpenAI has unveiled [GPT-6 Astra](https://www.testingcatalog.com/first-outputs-from-gpt-6-astra-model-from-openai/), calling it its most intelligent and aligned model yet. The first wave is reaching a limited set of organizations, with access expanding over the following days to ChatGPT Plus, Pro, Business, and Enterprise, the OpenAI API, Microsoft Azure, and Amazon Web Services Bedrock. Astra Pro is also coming to Pro, Business, and Enterprise plans, while enterprise administrators must enable access for their workspaces. > This is GPT-6 Astra. > > Anything you can do on a computer, Astra can do for you. Fast. [pic.twitter.com/gDd0IsewJw](https://t.co/gDd0IsewJw?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [September 3, 2026](https://x.com/OpenAI/status/2095595741528125780?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Astra targets computer use and professional work across browsers, email, calendars, CRM systems, documents, spreadsheets, presentations, software, scientific analysis, websites, and games. OpenAI reports 72.6% on OSWorld 2.0 at about 40 minutes per task, compared with 65.7% and about 75 minutes for GPT-5.6 Sol. An updated Codex harness helped Astra complete Mind2Web work 1.9 times faster than the current Sol experience. > Astra is very good at 3D modeling, and I can't wait for all of you to experience it, for now here is a little walkthrough on how I built the demo house for our launch blog post. From a Blender scene to a Unreal Engine 5 walkable experience. 🧵 [pic.twitter.com/xGqe5iXoRv](https://t.co/xGqe5iXoRv?ref=testingcatalog.com) > > — Thomas Ricouard (@Dimillian) [September 3, 2026](https://x.com/Dimillian/status/2095596700815516004?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) OpenAI is positioning Astra as its strongest software engineering model, reporting 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1\. A new experimental Codex context feature can retain notes across context windows and search earlier requirements and test results. Astra can continue working while asking asynchronous questions, make sensible assumptions, and pause when consequential choices require user input. > OPENAI 🔥: ChatGPT paid users will get one banked reset for every day they don’t have Astra access, starting today. > > \> Sam Altman - i am hopeful that you can use it this weekend! but can’t promise yet. > > This time I need [@patience\_cave](https://x.com/patience%5Fcave?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) to hold the ground for real, can we get at… [https://t.co/RUYghCiacE](https://t.co/RUYghCiacE?ref=testingcatalog.com) [pic.twitter.com/nhmQQEOsGf](https://t.co/nhmQQEOsGf?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [September 4, 2026](https://x.com/testingcatalog/status/2095777247735042076?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Cybersecurity is the sharpest edge of the release. Astra reached OpenAI’s Critical capability threshold and scored 100% on ExploitBench without production safeguards. It found and used two previously unknown zero-day vulnerabilities during evaluation, which OpenAI says it is disclosing to maintainers. The launch model will refuse advanced requests such as creating proof-of-concept exploits, while OpenAI Daybreak is set to widen access for defensive work. In one impossible-task evaluation, Astra never went beyond its authorized target, though its written reasoning proved harder to monitor than GPT-5.6 Sol. > GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex. It's also live in the API. > > It might take a few days to roll out to our Plus and Business users. Thank you for your patience. > > — OpenAI (@OpenAI) [September 4, 2026](https://x.com/OpenAI/status/2095968413646737608?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For OpenAI, Astra connects ChatGPT, Codex, its API business, and distribution through two major cloud platforms in one release. The API model is gpt-6-astra, priced at $10 per million input tokens and $50 per million output tokens. Fast mode offers up to twice the speed at twice the price. The company frames Astra as a broad step across knowledge work, coding, science, computer use, and security. [Source](https://openai.com/index/gpt-6-astra/?ref=testingcatalog.com) ### Lindy can now schedule meetings from email thread URL: https://www.testingcatalog.com/lindy-can-now-schedule-meetings-from-email-thread/ Last updated: 2026-09-03T16:30:28.000Z Lindy is an assistant that works inside email and calendar, and it has added a way to schedule meetings from the CC line. Adding the Lindy address to any thread where two people are trying to find a time hands off the coordination. Lindy reads what's already been said, pulls out the scheduling details, and fills in whatever is missing with sensible defaults. From there, the behavior splits two ways. If the thread has already settled on a time, Lindy books it. If no time has been agreed, it replies once with three open slots as clickable links. The meeting lands on the calendar, and the invites go out. Everything the person on the other end has to do stays inside the email they were already reading, with no account or scheduling page required on their side. Constraints written in ordinary language in the thread are read and respected, so a line asking for mornings only or a thirty-minute call carries through without a settings screen. Two conditions have to be true before anything happens. The Lindy address has to sit on the CC line, and the thread has to show that a time is actually being arranged. Where either is missing, the assistant stays out of the conversation. Scheduling by CC has been possible in Lindy before this, but as a workflow that had to be assembled first. The published template walks through copying a dedicated Lindy mail address out of an email trigger, setting a condition that decides whether a time was proposed or still needs searching, configuring a calendar action with working hours and daily meeting caps, and writing the reply text the assistant sends back. This release reaches the same outcome without that build, moving scheduling from something configured in advance to behavior that runs on the thread itself. SPONSORED Check out Lindy! [Learn more ](https://www.lindy.ai/?ref=testingcatalog.com) Lindy was founded in 2023 by Flo Crivello, previously a product manager at Uber and founder of the virtual office company Teamflow. The assistant already runs across several surfaces. It works in Gmail through a Chrome extension, answers inside Slack channels and direct messages, runs from iMessage, and joins calls on Meet, Zoom, and Teams to record and file notes into folders described in plain language. It connects to Gmail, Drive, Calendar, and Notion, plus more than a thousand other integrations; supports MCP servers; and keeps its memory in plain files that can be opened and edited. Anything carrying outside consequences waits on approval from a named person, the same guardrail that governs the scheduling flow. Plans start at 29.99 dollars per user per month, with SOC 2 Type II, GDPR, and PIPEDA compliance reported, and HIPAA available on the enterprise tier. ### Google releases Gemini 3.8 Flash and Flash Cyber URL: https://www.testingcatalog.com/google-releases-gemini-3-8-flash-and-flash-cyber/ Last updated: 2026-09-03T09:19:23.000Z Google has released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, splitting its new model family between agentic work and restricted cyber defense. The third Flash launch in six weeks arrives three weeks after 3.7 Flash. Both variants share foundational intelligence, and 3.8 Flash keeps the same introductory pricing: $0.75 per million input tokens and $3.75 per million output tokens. Gemini 3.8 Flash targets long-horizon coding, autonomous agents, and multi-step reasoning. Google says it beats most larger frontier models on DeepSWE v1.1 and scores 54.9% on HLE-Verified, with further gains on finance and legal agent benchmarks. Google attributes those results to additional reasoning steps and iterative tool calls. That can use more tokens at higher effort levels, so developers can select lower effort or stay on supported 3.7 Flash when compute efficiency matters. > Two new Gemini models are here to help scale your AI agents and secure code: > > 🔘 3.8 Flash: our most intelligent model yet with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. > > 🔘 3.8 Flash Cyber: our most capable… [pic.twitter.com/EEJDIMhRwp](https://t.co/EEJDIMhRwp?ref=testingcatalog.com) > > — Google DeepMind (@GoogleDeepMind) [September 2, 2026](https://x.com/GoogleDeepMind/status/2095175498967949359?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model is available now through the Gemini API in Google AI Studio and Android Studio, plus Google Antigravity and Stitch. Enterprises can access it in Gemini Enterprise. Google AI Pro and Ultra subscribers can use it in the Gemini app, AI Mode in Search, and Gemini in Google Sheets. > Gemini 3.8 Flash on DeepSWE 1.1, scores 73.7%! [pic.twitter.com/JW7fhVy4He](https://t.co/JW7fhVy4He?ref=testingcatalog.com) > > — Logan Kilpatrick (@OfficialLoganK) [September 2, 2026](https://x.com/OfficialLoganK/status/2095178478505328918?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Gemini 3.8 Flash Cyber focuses on vulnerability discovery and automated patching. Google reports a success rate above 70% on an internal benchmark covering complex codebases in 20 languages. Its 47.2% pass@1 result on CWE-Bench trails a leading frontier model’s 47.8% at lower cost. Chrome’s security team says it produced 2.6 times as many correct patches as the best commercial models tested. Wiz measured higher recall at lower cost, while Google’s Cloud Vulnerability Research team used it to find a critical foundational flaw in under two hours. [Google](https://www.testingcatalog.com/tag/google/) says both releases were accelerated by long-running agentic loops that recursively evaluate and refine the models, with training advances that included cybersecurity work. Google limits cyber access to trusted government authorities, critical-infrastructure operators, and software maintainers through its new Fairwind Program. The Cyber variant uses more permissive cybersecurity mitigations, while the main Flash model carries safeguards for chemical, biological, radiological, nuclear, and cyber-offense risks under Google’s Frontier Safety Framework. The company also reports stronger prompt-injection robustness in Gray Swan testing. [Source](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/?utm%5Fsource=x&utm%5Fmedium=social&utm%5Fcampaign=&utm%5Fcontent=) ### ICYMI: OpenAI's Astra crosses Critical cybersecurity threshold URL: https://www.testingcatalog.com/icymi-openais-astra-crosses-critical-cybersecurity-threshold/ Last updated: 2026-09-03T08:29:18.000Z OpenAI says Astra has crossed the Critical cybersecurity capability threshold under its Preparedness Framework, becoming the company’s first model at that level. With the right tools and access, Astra can find unknown flaws and develop working exploits across hardened systems without a person directing every step. OpenAI plans to release it soon, but its strongest cyber capabilities will first go to a small group of alpha testers, with Daybreak Blue access following for wider defensive use. > OPENAI 🔥: Astra will be "available soon," but its cybersecurity capabilities will be limited. > > \> Astra scored 100% on ExploitBench. > \> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.… [https://t.co/xKz3xloRnV](https://t.co/xKz3xloRnV?ref=testingcatalog.com) [pic.twitter.com/Z63iY3SHE9](https://t.co/Z63iY3SHE9?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [September 1, 2026](https://x.com/testingcatalog/status/2094889418041516441?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) OpenAI combined public and private benchmarks with expert-led tests. Astra scored 100% on ExploitBench and was substantially more capable and token-efficient than GPT-5.6 Sol in vulnerability discovery and exploit development. On an internal set of 20 recently disclosed high-severity V8 flaws, Astra delivered higher arbitrary-code-execution rates with far fewer output tokens. It also found two zero-days and used them in an exploit chain now being disclosed to maintainers. The results reflect Daybreak Blue access, not the default production setup. In hands-on assessments, Astra built a browser compromise chain that escaped the sandbox and ran host commands when an HTML file was opened. It also combined flaws in a hardened operating system to escalate an unprivileged user to root. OpenAI says these results meet its Critical bar for zero-day exploitation across hardened systems or end-to-end attacks from a high-level goal. OpenAI delayed parts of Astra’s training and release while adding refusals, safety classifiers, offline detection, threat disruption, cross-conversation context handling, and monitoring that can stop unauthorized activity. Astra refused 91.5% of requests in cyber jailbreak evaluations, versus 59% for GPT-5.6 Sol. In honeypot tests based on the Hugging Face incident, Astra did not target surrounding systems or try to bypass auto-review denials. For [OpenAI](https://www.testingcatalog.com/tag/openai/), Astra is an early test of how its Preparedness Framework governs models capable of consequential work. The company says advanced access will expand cautiously as red-teaming, regression testing, and 24/7 response continue. Chain-of-thought monitoring will examine reasoning and actions for unauthorized behavior, while legitimate tasks may be slowed, paused, or stopped in ChatGPT, Codex, or the API. A system card at launch will provide more detail on Astra’s safety, security, alignment, and evaluations. [Source](https://openai.com/index/path-to-astra/?ref=testingcatalog.com) ### Wonderful raises $550M for its enterprise AI agent platform URL: https://www.testingcatalog.com/wonderful-raises-550m-for-its-enterprise-ai-agent-platform/ Last updated: 2026-09-02T14:45:17.000Z Wonderful has raised a $550 million Series C at a $5 billion valuation, led by Insight Partners. The company builds and runs AI agents for large enterprises, covering customer conversations across voice, chat, email, and in-app channels, employee requests that need action inside internal systems, and back-office work such as reconciling data and moving processes forward without a person involved. The platform is paired with full-stack engineering teams that sit inside customer environments in the market they serve, and that pairing is where the money is going. > Wonderful has raised a $550M Series C at a $5 billion valuation. > > This latest funding was led by [@insightpartners](https://x.com/insightpartners?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), with [@salesforce](https://x.com/salesforce?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) joining as new investors and [@IndexVentures](https://x.com/IndexVentures?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), [@BessemerVP](https://x.com/BessemerVP?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), [@IVP](https://x.com/IVP?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), [@VineVenturesLP](https://x.com/VineVenturesLP?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), and [@9yardscapital](https://x.com/9yardscapital?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) returning. > > In 20 months since launch, we're in… [pic.twitter.com/2LIUmMgu20](https://t.co/2LIUmMgu20?ref=testingcatalog.com) > > — Wonderful (@wonderful\_ai) [September 2, 2026](https://x.com/wonderful%5Fai/status/2095125312929464388?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Teams build agents in plain language or in code, attach reusable skills, connect systems of record over API or MCP, and test variants side by side before shipping. Guardrails set what an agent can and cannot do, and full reasoning traces are recorded for every action, so an audit shows what an agent did and why. Live dashboards cover usage, efficiency, and outcomes. In January, the company shipped Agent Builder, an autonomous agent that builds, tests, and refines other agents, running on Anthropic's Claude. Wonderful's stated position is that the hard part of enterprise AI is delivery. Agents that pass a demo still need integration into core systems, localization, compliance review, and continuous tuning, and enterprises pay for that gap in months of pilots that never carry live traffic. The company answers it by placing local teams inside customer environments, co-building with in-house staff and tuning agents to the language, regulation and operating context of each country. Wonderful says this moves deployments from pilot to production in days and weeks. ELTA Hellenic Post reports scaling agentic customer support by close to four times within two months of go-live, holding an 86 percent success rate in production. SPONSORED Test out Wonderful [Learn more ](https://www.wonderful.ai/?ref=testingcatalog.com) Wonderful was founded in early 2025 by Bar Winkler and Roey Lalazar and came out of stealth that July. Twenty months on, it reports hundreds of agents, systems, and workflows running in production across a dozen verticals and more than 30 markets, including telecom, financial services, healthcare, and manufacturing. The round brings Salesforce in as a new investor, with Index Ventures, Bessemer, IVP, Vine Ventures and 9Yards returning. The capital is earmarked for customer delivery, hiring engineers and forward-deployed staff, and widening the product across the whole enterprise. ### Catch launches an AI assistant with phone calling support URL: https://www.testingcatalog.com/catch-launches-an-ai-assistant-with-phone-calling-support/ Last updated: 2026-09-02T13:50:59.000Z Catch has launched an AI administrative assistant that works the inbox, runs the calendar, and places phone calls on behalf of executives, announcing a seed round alongside the release. It costs $ 99 a month, includes a 10-day trial, and signs in through a Google or Microsoft account. The category it enters is crowded with tools that stop one step short. They flag a double booking, draft a reply, or produce a morning brief, then return the work to the person who wanted it gone. Catch is built to close the loop. It emails the other side to move a meeting, sends the reply once intent has been confirmed in a sentence, and places outbound calls in the account holder's name. > Big NEWS: Today we're launching Catch AI, an admin-assistant for busy executives. > > We started Catch with a simple idea: > AI should actually do the administrative work executives hate doing themselves. > > Not suggest. Not summarize. Do. > \>> [pic.twitter.com/HI80mrmpay](https://t.co/HI80mrmpay?ref=testingcatalog.com) > > — Nir Sabato (@therealnirs) [September 2, 2026](https://x.com/therealnirs/status/2095134650431586686?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Voice is where the product separates from the field. Catch dials restaurants, clinics, hotels, airlines, and suppliers, navigates automated menus until a person picks up, and delivers a summary and full transcript once the call ends. It speaks in its own voice, identifies itself as software calling on someone's behalf, and dials only on instruction. The line runs inbound too, with Catch ringing through a hands-free walkthrough of the day's schedule, taking decisions out loud, then acting on them after the hangup. There is no app to install. Catch operates through Slack, WhatsApp, iMessage, email, and phone, and connects to Google Calendar and Outlook, Zoom and Google Meet, Notion, HubSpot, Asana, ClickUp, Todoist, and meeting recorders including Granola, Otter, and Fireflies. A free scheduling concierge named Uta sits outside the subscription entirely, copied onto any email thread without an account, working with each participant separately by email or text before returning the calendar invite. Business travel is the least finished piece. It runs as a beta for selected customers, shortlisting flights and hotels against a stated budget and meeting schedule, rechecking fares and cancellation terms, and booking only after explicit approval. Fares requiring passport or identity documents are unsupported. The company reports SOC 2 Type II compliance, CASA Tier-2 certification and Google verification, and counts over 5,000 meetings scheduled, 20,000 tasks delegated and 200,000 emails handled to date. SPONSORED Test out Catch for yourself [Learn more ](https://www.catchagent.ai/?ref=testingcatalog.com) Catch was founded by Nir Sabato and Yoav Ramon and runs on the company's own AI stack. Sabato, the chief executive, was an investor at Entree Capital before starting the company and previously led strategy at Fiverr and BLEND. Ramon, the chief technology officer, has spent over a decade shipping AI into production, building speech and language systems at Hi Auto and then running large language models in regulated healthcare as chief technology officer of Nym. Neither the size of the seed round nor its lead investor has been disclosed. ### Muse superapp from Meta and Ava model with computer use URL: https://www.testingcatalog.com/muse-superapp-from-meta-and-ava-model-with-computer-use/ Last updated: 2026-09-02T11:50:26.000Z Meta is moving closer to launching its planned agent super app, previously known internally as Project Hatch. TestingCatalog has found that the product is being prepared under the launch name Muse, bringing it under the same branding Meta already uses for its new generation of AI models. > META 🔥: Project Hatch will be released under the name “Muse” and will arrive with a waitlist! > > \> Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post below. > > Joined 👀👀👀 [https://t.co/TQqXDYqk4b](https://t.co/TQqXDYqk4b?ref=testingcatalog.com) [pic.twitter.com/VX8YacN3G1](https://t.co/VX8YacN3G1?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [September 1, 2026](https://x.com/testingcatalog/status/2094827945843921332?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The clearest change is on iOS. An app previously used for internal testing has transitioned into a Muse waitlist experience. Joining the waitlist appears technically possible, although Meta has not opened it publicly yet. 💡 We plan to share access details for iOS users in this Sunday's TestingCatalog newsletter. Make sure you've subscribed! ## Sign up for TestingCatalog AI News TestingCatalog - Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors Subscribe Email sent! Check your inbox to complete your signup. No spam. Unsubscribe anytime. Meta’s existing desktop app is also changing. Recent updates add smaller additions such as theme support, secondary color selection, and a monochrome theme. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/Screenshot-2026-09-01-at-21.35.08.webp) More importantly, **Meta has added a setting for computer use** and appears to be **testing a model variant called Ava**, described as supporting computer control. Ava itself is not currently accessible. These additions provide more context around what Muse could become. Computer use would allow the agent to operate desktop applications as part of longer tasks, while browser control also appears to be under development. This would put Muse closer to products such as Codex and Claude that are increasingly built around autonomous, multi-step work rather than conventional assistant chats. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/09/Screenshot-2026-09-01-at-21.44.49.webp) The direction also matches Meta’s broader model strategy. Muse Spark 1.1 and 1.2 already emphasize tool use, coding, and computer control, while Meta has publicly framed its AI roadmap around agents that can pursue goals and take action. Reports about [Hatch](https://www.testingcatalog.com/exclusive-deeper-look-into-hatch-agent-from-meta/) have previously described a product capable of using websites, managing schedules, sending emails, and building custom tools, with Meta reportedly considering premium pricing as high as $200 per month. The waitlist transition may also mean a more controlled rollout than previously expected. Rather than opening Muse broadly at launch, Meta could announce it first and gradually admit users while computer and browser control mature. ### Anthropic launches Claude Fable 5.1 and Mythos 5.1 URL: https://www.testingcatalog.com/anthropic-launches-claude-fable-5-1-and-mythos-5-1/ Last updated: 2026-09-01T23:59:59.000Z Anthropic has launched Claude Fable 5.1 and Claude Mythos 5.1, the same model with different safeguards for coding, knowledge work, and scientific research. Fable is generally available now, while Mythos is limited to vetted cyber defenders and life scientists. Developers can call claude-fable-5-1 via the Claude API and access Fable through Anthropic's products, AWS, Google Cloud, and Microsoft Azure. > We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. > > They're the world’s most advanced models for coding and knowledge work. [pic.twitter.com/8P9PSrWPi3](https://t.co/8P9PSrWPi3?ref=testingcatalog.com) > > — Claude (@claudeai) [September 1, 2026](https://x.com/claudeai/status/2094848572143407483?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Cache reads now cost 75% less at $0.25 per million tokens. Anthropic estimates typical costs will fall 25%, with savings on highly agentic work reaching about 45%. Other rates remain $10 per million input tokens and $50 per million output tokens. Fable defaults to High effort in Claude Code and Medium effort in Claude Cowork and Claude. Anthropic also introduced Enterprise Frontier Safeguards, which keep customer data inside infrastructure controlled by the customer while retaining misuse protections. It will roll out to enterprises in phases later this fall. Until then, eligible customers can use Fable with zero data retention. Anthropic reports 60% fewer false-positive cyber interventions per Claude Code session and 85% fewer biology interventions on benign elementary questions than its earlier safeguards. > As well as being capable of much higher performance than Fable 5, it can also achieve similar or better results at a much lower cost when set to lower effort levels. [pic.twitter.com/xc7sBvEx5Y](https://t.co/xc7sBvEx5Y?ref=testingcatalog.com) > > — Claude (@claudeai) [September 1, 2026](https://x.com/claudeai/status/2094848585451847721?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Anthropic reports Fable scored 52.6% on Terminal-Bench-Science 0.1 at maximum effort, versus 24.7% for Fable 5, and 73.4% on CursorBench 3.2.0\. Mythos reached 60.9% on Terminal-Bench 4.0, compared with 55.8% for Fable. Production safeguards were active and could lower some scores. In research tests, Mythos designed protein binders sent for outside validation, with Anthropic reporting a hit rate near 50% across 12 targets. Fable trained a neural network to produce a Venus elevation map at two-to-three-kilometer resolution, up from 10 to 20 kilometers, while Mythos wrote GPU kernels for seven open-source biology models with speedups up to 2.5 times. The map will be released under Creative Commons, and Anthropic plans to open-source the biology optimizations. Anthropic tested both models for chemical and biological risk, cybersecurity, agentic safety, alignment, prompt injection, and reward hacking. Mythos access runs through the Cyber Verification and Life Sciences Verification programs and powers Claude Security. New-model outputs carry an invisible numerical watermark for EU AI Act transparency, backed by a private-preview detection API for eligible groups. [Source](https://www.google.com/url?q=https://www.anthropic.com/claude-fable-and-mythos-5-1&source=gmail&ust=1788377822273000&usg=AOvVaw2G71MXLOkY1NP%5FIk%5FCi8Mx) ### Visko launches Orbis, a new real-time AI video model URL: https://www.testingcatalog.com/visko-launches-orbis-a-new-real-time-ai-video-model/ Last updated: 2026-09-01T16:24:24.000Z Visko has launched Orbis, a video model that generates a scene in real time and keeps generating it while the prompt changes underneath. Conventional video tools take a prompt, run for a while, and hand back a finished clip that is fixed once it arrives, so any change means paying for a fresh generation and discarding what the last one built. Orbis runs continuously instead, streaming frames as they are produced, and a new instruction typed mid-stream lands in the output without restarting the run. 0:00 /0:21 1× Visko calls this class of system a Live Model, meaning a model that runs as an ongoing process tied to real time and carries its state forward for as long as it runs. A memory layer keeps subjects, scenes, and style consistent as the video extends, and because that memory sits inside the model instead of a growing transcript, the cost of carrying the past stays roughly flat as a run gets longer. The company reports hour-long runs that hold steady on visual quality and color without drifting. Orbis accepts text, an image, or an existing video as a starting point, takes prompts in multiple languages, and delivers 4K at 24 frames per second through a serving stack built for continuous generation. Visko puts the delay between a typed change and a visible one under a second on average. Internally, it generates at a lower resolution and passes the result through a streaming upscaler, so 4K arrives without waiting for a full sequence to finish. The technical report measures Orbis on 74 test cases, each a video of one to three minutes carrying six prompt changes, with every segment scored against whichever prompt was active at the time. On that set, Orbis takes the top position for perceptual quality, motion quality, and prompt alignment under the DOVER and VideoAlign metrics, measured against real-time video systems including Self-Forcing, LongLive, and SANA-Video. On a separate aesthetic measure, it places third, behind SANA-Video and LongLive. Orbis runs on the Visko site behind a sign-in, with demos covering branching stories, voice-driven generation, robot training scenes, and a freeform playground for building a world from text or a reference image. Two generation modes are offered, one for high-motion scenes and one for slower scenes with finer detail. Intended users span entertainment and live streaming, virtual companions, education, and robotics teams wanting a simulated environment to train in before a machine acts in the real world. SPONSORED Test out Visko Orbis for yourself [Learn more ](https://www.visko.ai/?ref=testingcatalog.com) Visko Platform is based in Sunnyvale, California, and was founded in 2025\. Orbis sits alongside two narrower models from the same team: Morphe, which swaps or animates a person while leaving the rest of the frame intact, and Kinesis, which adds or alters elements at a chosen point in space and time. The company argues video generation is moving from tools that return a clip toward processes that keep running, priced by the hour of a world staying live. The launch comes alongside a $10 million pre-seed round. ### Google develops AI Rooms for Gemini Enterprise URL: https://www.testingcatalog.com/google-develops-ai-rooms-for-gemini-enterprise/ Last updated: 2026-08-31T14:07:43.000Z Google is prototyping a new Gemini Enterprise feature called Rooms, which could turn the platform into a shared workspace where teams collaborate with both colleagues and Gemini around a specific objective. TestingCatalog found references to the feature in recent Gemini Enterprise builds, but it remains marked as a prototype, and Google has not yet indicated plans to ship it publicly. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Gemini-Enterprise-08-31-2026_12_52_AM.jpg) Rooms are described as an “all-in-one place” environment where Gemini becomes an expert using connected files while keeping team conversations focused. Creating one involves defining a goal, a playbook describing how Gemini should operate, and a knowledge base. Teams can then add members and files, maintain shared context, and communicate with each other or Gemini inside the same space. The concept appears closely related to [Projects](https://www.testingcatalog.com/google-expands-gemini-for-business-with-shareable-projects/), which Google publicly introduced at Cloud Next 2026 as shared workspaces for humans and agents. Projects already combine conversations with sources from services such as Google Workspace, Microsoft OneDrive, NotebookLM, and team chats. Rooms appear to take that idea further by giving the workspace an explicit objective, operating instructions, and a more structured knowledge layer. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Rooms-Gemini-Enterprise-08-31-2026_12_53_AM.jpg) For enterprise teams, this could position Gemini closer to project management and operational collaboration software rather than serving primarily as an assistant layered over company data. One particularly useful direction would be connecting context gathered during Google Meet calls with room decisions, documentation, tasks, and ongoing conversations, although there is currently no evidence that Meet integration is part of the prototype. Google has been steadily repositioning [Gemini](https://www.testingcatalog.com/tag/gemini/) Enterprise as an agentic workplace platform, with long-running agents, Projects, Canvas, an Agent Gallery, enterprise connectors, and persistent agent memory all forming parts of that strategy. Recent interface changes point in the same direction: the Customization area is being reorganized to include dedicated Agents and Memory sections, with agents potentially moving into a list where users can inspect and configure them. For now, Rooms remain an internal prototype. Their final name, availability, supported data sources, and whether they will replace or coexist with Projects are still unclear. ### Perplexity prepares Hybrid mode for Computer on Mac URL: https://www.testingcatalog.com/perplexity-prepares-hybrid-mode-for-computer-on-mac/ Last updated: 2026-08-31T08:48:58.000Z Perplexity is preparing to bring local inference to a much wider group of Mac users through a new Hybrid Mode for Perplexity Computer. The company previously announced hybrid agentic inference in June and said support would extend beyond Windows to Macs and Linux. More recently, it launched Portable Computer for NVIDIA DGX Spark, where Computer can run entirely on local hardware. > Today we’re launching Portable Computer on [@NVIDIA](https://x.com/nvidia?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) DGX Spark. > > Portable Computer is a fully local version of Perplexity Computer, where the entire runtime: orchestrator LLM, subagent LLM, agent harness all run on your local hardware. No cloud dependency. [pic.twitter.com/plVWz5PaAw](https://t.co/plVWz5PaAw?ref=testingcatalog.com) > > — Perplexity (@perplexity\_ai) [August 25, 2026](https://x.com/perplexity%5Fai/status/2092268362386780270?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) On Mac, the implementation appears different. Hybrid Mode lets Computer continue orchestrating in the cloud while automatically delegating suitable subtasks to a model running locally. Cloud models handle heavier reasoning, while local inference handles lighter work without consuming Computer credits and keeps the associated data on the Mac. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-30-at-01.00.39.png) ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-30-at-00.59.13.png) Users can enable Hybrid Mode from the Computer model selector and configure local inference if no model is installed. Three downloadable options are currently being prepared: a recommended Perplexity model around 19 GB requiring at least 32 GB of memory, a roughly 17.4 GB Qwen 32B option with similar requirements, and a smaller Gemma-based model around 5.6 GB that can run on 16 GB Macs. The exact identity of Perplexity's own model remains unclear. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-31-at-00.45.36.png) Another feature, Privacy Gate, adds a separate local model that inspects data before it is sent to the cloud. If it detects personal or sensitive information, Computer can ask the user how to handle it before transmission. This could be particularly relevant to users and companies operating under GDPR requirements in the EEA. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-31-at-00.45.05.png) Computer is evolving from a cloud-only agent into an orchestration layer spanning cloud infrastructure and user-controlled hardware. The company has already publicly described this architecture as hybrid agentic inference, positioning local processing around privacy and lower inference costs while retaining access to frontier models. The Mac controls described here remain hidden, and [Perplexity](https://www.testingcatalog.com/tag/perplexity/) has not announced their release date or whether access will be limited by subscription tier. A fully local Mac orchestrator, similar to Portable Computer on DGX Spark, also does not appear to be confirmed yet. ### Google prepares Interactive Reports for Gemini Notebook URL: https://www.testingcatalog.com/google-prepares-interactive-reports-for-gemini-notebook/ Last updated: 2026-08-30T07:51:09.000Z Google is preparing another major addition to Gemini Notebook called Interactive Reports. A new notice has appeared in the Studio panel stating, “New: you can now create interactive reports,” with a button that opens the existing Create Report customization flow. This type of release messaging suggests Google is already preparing the feature for users, potentially within the coming weeks. Inside the report creator, a new “Interactive” option is marked as New and described as “An interactive report with embedded Studio content.” Another preset, “Overview,” can create an interactive summary of key information while incorporating Studio content directly into the result. Users can select a language and provide detailed instructions describing the report they want. Google’s example asks for a formal competitive review of the 2026 functional beverage market, covering competitors, distribution, and pricing to support a launch strategy. This points toward research-heavy use cases where Gemini Notebook combines multiple sources into something closer to a structured research workspace than a static document. ![Gemini Notebook](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Gemini-Notebook-08-30-2026_12_12_AM.jpg) The exact types of Studio content that can be embedded remain unclear. Previous builds also contained placeholders for [Canvas](https://www.testingcatalog.com/google-tests-canvas-and-connectors-on-notebooklm/) and [Web Page](https://www.testingcatalog.com/google-is-working-on-interactive-apps-for-gemini-notebook/) outputs, resembling the artifact-style experiences available elsewhere in Gemini. Those concepts could remain separate, or some of that work may now be converging into Interactive Reports. The feature fits Google’s broader direction for Gemini Notebook. In July, Google expanded Studio beyond traditional summaries with support for generated documents, charts, spreadsheets, presentations and other editable outputs, while positioning Notebook as a research environment capable of handling more complex projects. Interactive Reports would push that strategy further by combining generated research with other Studio outputs inside one report. If the current release notice reflects Google’s rollout plans, this may be one of the next [Gemini Notebook](https://www.testingcatalog.com/tag/notebooklm/) features to surface publicly. ### First outputs from GPT-6 "Astra" model from OpenAI URL: https://www.testingcatalog.com/first-outputs-from-gpt-6-astra-model-from-openai/ Last updated: 2026-08-30T07:39:48.000Z OpenAI appears to be moving Astra closer to release, with reports pointing to expanded internal testing and a new checkpoint identified as “mozaik-alpha-fdm.” The latest surfaced outputs were reportedly generated zero-shot with Max effort, where Astra spends substantially longer reasoning than GPT-5.6 Sol. 0:00 /0:38 1× Voxel Castle output from Astra The examples point to a strong focus on coding and visual software creation. Astra reportedly produced a GTA 2-style game in one attempt, alongside detailed websites, 3D objects, and voxel environments. The main difference visible in these samples is attention to small implementation details and the amount of complete functionality produced from a single prompt. Developers, designers, and founders building prototypes could be among the main beneficiaries if this performance carries over to the released model. > 🚨 Major Scoop: > > OpenAI has just expanded internal testing of GPT Astra, codenamed "mozaik-alpha-fdm", signifying release might be near > > AND of course like previous times, we have the FIRST EVER public outputs of it for y'all😉 > > Both outputs are zero-shot on Max effort. Frontend… [https://t.co/PxvM8eIQqp](https://t.co/PxvM8eIQqp?ref=testingcatalog.com) [pic.twitter.com/geZcl3e6tl](https://t.co/geZcl3e6tl?ref=testingcatalog.com) > > — Lentils (@Lentils80) [August 29, 2026](https://x.com/Lentils80/status/2093617080327127456?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Astra could arrive through [ChatGPT](https://www.testingcatalog.com/tag/chatgpt/), Codex or both, especially given OpenAI's focus on agentic coding and long-running work. OpenAI has not announced a release date or confirmed that “mozaik-alpha-fdm” is an official codename. Separate reports have suggested a launch within weeks, while the newly expanded testing provides another indication that preparations are progressing. > 🚨 Astra < left > vs Fable 5.1 < Right > > > Prompt : generate a svg of ana de armas as detailed as possible [https://t.co/6LacRvPYoW](https://t.co/6LacRvPYoW?ref=testingcatalog.com) [pic.twitter.com/QMrNMvKviu](https://t.co/QMrNMvKviu?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [August 29, 2026](https://x.com/chetaslua/status/2093663148528341202?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) One major variable remains timing: safety approval. OpenAI has officially acknowledged Astra as an upcoming model and said its internal evaluations showed major advances in agentic coding and cybersecurity, potentially reaching its Critical cyber capability threshold. Some Astra workloads were subsequently paused while stricter safeguards were introduced. OpenAI has also confirmed that relevant government agencies and selected AI safety organizations will participate in testing. > \[GPT Astra\] Pagoda Voxel > \~56k tokens; \~38 minutes [pic.twitter.com/qsmx24ZyjA](https://t.co/qsmx24ZyjA?ref=testingcatalog.com) > > — lyra (@lyraxana) [August 30, 2026](https://x.com/lyraxana/status/2093882141750763982?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > 还有这个 [pic.twitter.com/T0WNrNXLOV](https://t.co/T0WNrNXLOV?ref=testingcatalog.com) > > — Rosenda Garica (@GaricaRosen6779) [August 29, 2026](https://x.com/GaricaRosen6779/status/2093712275475693802?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) This makes government evaluation part of the actual deployment process rather than simply another rumor. It also means a technically ready Astra may still remain gated until OpenAI is satisfied with its monitoring, alignment, and containment systems. > Astra may be the first time OpenAI have ever underhyped a model > > The gap between step changes in LLM capability keeps shrinking. And things are going to look so different 3 months, 6 months, a year from now, nevermind 5. > > September is going to be a wild ride > > — leo 🐾 (@synthwavedd) [August 29, 2026](https://x.com/synthwavedd/status/2093488804602487177?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Meanwhile, several people with histories of early access to frontier models have hinted that Astra represents a much larger capability jump than recent [OpenAI](https://www.testingcatalog.com/tag/openai/) releases. GPT-6 remains a plausible public name, but OpenAI has given no indication that it has chosen this branding. Source: [zAI Discord](https://discord.gg/3Km9qnH6D?ref=testingcatalog.com) ### Tencent released open-source Hy4 preview model URL: https://www.testingcatalog.com/tencent-released-open-source-hy4-preview-model/ Last updated: 2026-08-29T12:39:08.000Z Tencent has released Hy4 preview, a 770-billion-parameter flagship model with 49 billion active parameters and a one-million-token context window. The company calls it its most capable model so far and says it delivers the largest generation-over-generation gain it has measured. The open-source release targets long-horizon software engineering, document-heavy office work, game development and scientific research. > We compressed Hy4-preview from 1.5TB to ~200GiB GGUF and it still works well ! > > Meet MIX-STQ1\_0.The trick isn’t just going low, it’s deciding where: calibration data picks each layer’s bit-width, some down to 1.31-bit STQ1\_0, some up to 2.06-bit IQ2\_XXS. Same budget, lower… [https://t.co/HItkQ1TSxA](https://t.co/HItkQ1TSxA?ref=testingcatalog.com) [pic.twitter.com/xK8d3BH5br](https://t.co/xK8d3BH5br?ref=testingcatalog.com) > > — Tencent Hy (@TencentHunyuan) [August 29, 2026](https://x.com/TencentHunyuan/status/2093572224342954019?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For software projects, Hy4 preview is designed to understand, plan, debug and verify work across extended sessions, with added focus on the visual quality of front-end output. It can also turn scattered files into documents, spreadsheets and presentations, including work involving data analysis, equations and financial models. Tencent says the model can build playable prototypes from one prompt and continue refining complex projects in game engines over multiple turns. Its research scope includes AI, condensed matter physics and pure mathematics. In one test, Hy4 preview managed several Codex sessions in parallel and changed its research direction as results arrived. On a small-model post-training task, it coordinated work across several evaluation targets and beat Codex working independently on all eight benchmarks, according to Tencent. The company presents this as evidence that the model can judge research direction, organize experiments, and iterate on complex R&D work. 0:00 /0:11 1× A blind internal comparison involved 163 Tencent experts rating outputs across 203 engineering tasks. Hy4 preview averaged 2.99, slightly above GLM 5.3 (2.92) and Kimi K3 (2.94). Against GLM 5.3, it recorded 46.8% wins, 12.8% ties, and 40.4% losses. Against Kimi K3, it posted 51.2% wins, 7.9% ties and 40.9% losses. Tencent built Hy4 preview with training data shaped around work performed by its software engineers, game developers, finance analysts, and security experts. It continues to co-design the model with CodeBuddy and WorkBuddy, while access is available through Tencent Cloud, OpenRouter, Yuanbao, ima, and public repositories. API pricing per million tokens is $0.042 for cached input, $0.834 for input, and $2.501 for output. Tencent describes Hy4 preview as an early version and acknowledges that it can reason longer than needed on complex tasks and over-verify its own work. [Source](https://hy.tencent.ai/research/hy4-preview?ref=testingcatalog.com) ### Exclusive: Deeper look into Hatch Agent from Meta URL: https://www.testingcatalog.com/exclusive-deeper-look-into-hatch-agent-from-meta/ Last updated: 2026-08-28T20:45:20.000Z ## What was reported before Meta is preparing to launch Hatch, an autonomous consumer agent designed to perform online tasks rather than only answer questions. Business Insider reported that Meta has expanded internal testing beyond its Superintelligence Labs and described Hatch as a personal agent that continues working on a user’s behalf even when its app is closed. [Previously disclosed examples](https://www.businessinsider.com/meta-hatch-personal-ai-agent-capabilities-employees-memo-2026-8?ref=testingcatalog.com) include filling out forms, making purchases, conducting deep research, ordering food, booking restaurants, and finding a dog sitter. Hatch can connect directly with email, calendars, Spotify, Instagram, and OpenTable. It is also expected to retain more personal context than typical assistants while asking users to approve sensitive actions before carrying them out. Users will reportedly be able to name their agent, shape how it speaks, and decide what it should pay attention to. One employee created an agent called Veda that helped find a dog sitter, purchase a Father’s Day gift, and change its owner’s sleep habits. Separate reporting indicated that Hatch could launch in the coming weeks. Meta has considered subscription tiers, including a premium plan costing as much as $199.99 per month. Early prototypes reportedly included a customizable dashboard where agents could create tools such as fitness trackers and travel itineraries. ![Project Hatch mention](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Meta-AI-08-27-2026_10_50_PM.jpg) Project Hatch mention ## What we found on top Unreleased Hatch materials reviewed for this story reveal a much broader personal workspace than previously reported. The service appears to organize a user’s life across dedicated areas for conversations, goals, ideas, updates, and a personal library. People may be able to create long-running projects, schedule recurring work, save results in separate spaces, and return to completed research or generated files later. The onboarding process has three main stages: connecting accounts, choosing initial tasks, and designing the agent’s identity. Hatch introduces these stages with the messages “Hatch is better with connectors,” “Put your agent to work,” and “Customize your agent.” Suggested tasks, names, personalities, and avatars appear to be tailored to each account. Users can select starter assignments, edit the agent’s name, choose how it behaves, and replace its avatar. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.08.58.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.09.36.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.10.58.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.11.07.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.11.24.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.11.50.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.12.00.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.12.13.png) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-28-at-20.12.30.png) Leaked Project Hatch source code Hatch also appears capable of opening websites, working with files, and producing shareable artifacts. People could ask it to find information online, read or revise documents, organize stored material, and present results in reusable formats. Larger assignments may be divided among additional agents, while shared-agent controls suggest that users could reuse or distribute specialized assistants. The product contains dedicated experiences for voice, media, task progress, browser work, and approvals. Users appear able to see what Hatch is doing, review completed steps, and intervene before consequential actions. This suggests a workspace where several assignments can remain active at once, rather than a single conversation that ends when the window closes. Privacy setup is more extensive than previously disclosed. Hatch can prepare an encrypted private environment protected by a recovery PIN. The PIN is intended to unlock personal agent data on another device, while warning screens prevent people from continuing normally when the service cannot verify that their environment is protected. Mobile onboarding explicitly references Apple TestFlight and the Google Play Store, indicating planned iPhone and Android access alongside the web product. The service also contains invite-code, waitlist, age-check, disclosure, and NDA flows, suggesting Meta can control access by account or testing group before a wider rollout. These findings come from pre-release product materials. Some capabilities may remain restricted, change before launch, or arrive in later phases. ## What does all this mean These findings suggest Project Hatch has been tested as a **standalone Meta product** rather than merely a feature inside Meta AI. **Separate web, iOS, and Android implementation**s appear to exist, although Meta probably would still connect Hatch closely with its recently released macOS app. The iOS TestFlight build has already reached version 4, pointing to several internal iteration cycles. > BREAKING 🔥: "Project Hatch" from Meta will be a superapp with browser & computer use, /goals support, persistent cloud environments, spaces, and more! > > Here is what we know so far 👀 > \> Hatch has been tested internally as a standalone web, iOS, and Android app. > \> It supports… [pic.twitter.com/oE4fFfRy47](https://t.co/oE4fFfRy47?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 28, 2026](https://x.com/testingcatalog/status/2093439212502630667?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The scope appears unusually broad. Hatch is shaping up as Meta’s answer to "Superapps," with products such as Codex, Claude Desktop, and users potentially able to create and share multiple specialized agents. Current development references **desktop and browser control, long-running goals, memory, connectors, and Spaces** for separating work by project. Agents also appear designed to operate inside cloud VMs, potentially allowing **environments and tasks to persist** across devices. **Voice support** is referenced as well. Meta may not expose everything at launch. A first release could focus on browser and computer use before more advanced agent functionality arrives. Other work points to identity verification, **generated avatars for "hatched" agents**, and basic **chat themes**. More importantly, references to **Messenger Companion and WhatsApp Companion** suggest Hatch agents could eventually be reached directly through Meta’s messaging products. Native WhatsApp distribution would give Meta a route to bring autonomous agents to an audience few competitors can match. This direction fits [Meta’s](https://www.testingcatalog.com/tag/meta/) broader product strategy. Muse Spark was built around tool use, computer use, and agentic tasks, while Meta AI can already plan work, connect to email and calendars, and perform multi-step actions. Hatch appears to take that foundation further by turning it into a dedicated agent platform. Mentions of a [waitlist](https://www.testingcatalog.com/meta-prepares-hatch-agent-under-waitlist-and-social-media-skills/) suggest the initial rollout could be gated. September remains a plausible window, but it is still unclear whether desktop and mobile clients will arrive together. Source: Anonymous reporter ### Vellum Mobile brings Voice Mode and Mac control to iOS URL: https://www.testingcatalog.com/vellum-mobile-brings-voice-mode-and-mac-control-to-ios/ Last updated: 2026-08-27T18:03:57.000Z Vellum has taken its mobile app to full parity with its desktop and web clients, giving iPhone and Android owners the same assistant that until now did its heaviest work on a Mac. The app is free on the App Store and Google Play Store, and connects to either a cloud assistant or a self-hosted one. > Vellum is now available on iOS and Android. > > Life doesn’t happen behind your desk and your Personal AI shouldn’t be stuck there. > > Free, open-source, and has all the same features you already love: > > \> use with local or cloud assistants > \> conversational voice mode > \> 1-click… [pic.twitter.com/piutJTG33P](https://t.co/piutJTG33P?ref=testingcatalog.com) > > — vellum (@vellum\_ai) [August 27, 2026](https://x.com/vellum%5Fai/status/2093020764400136289?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Voice mode carries most of the release. Conversations run natively on the device and open from Siri, the Action Button, the Control Center, or a home screen shortcut. Once a session is running, Live Activities on the Lock Screen and in the Dynamic Island show what the assistant is doing and let users approve or deny sensitive actions without reopening the app. The camera can be turned on mid-call, so a document, a broken part, or a screen can be shown instead of described, and the assistant talks through the frame in real time. Speech recognition defaults to multilingual, with a picker for locking a single listening language when accuracy matters more than switching. 0:00 /0:06 1× The second half of the release is to reach back into the desktop. While a paired Mac is awake and running the desktop app, the assistant can write and edit local files, move files between its own workspace and the machine, run shell commands for deploys, scripts, and installs, and drive the browser and other applications. Recurring schedules can be created or changed from anywhere. App building, document drafting, note-taking, and image generation all run from the phone, and teams can pair self-hosted assistants by URL when running the open-source daemon on their own hardware. Vellum is built by Vocify Inc., a New York company founded in 2023 by Akash Sharma, Noa Flaherty, and Sidd Seethepalli. It went through Y Combinator's Winter 2023 batch and has raised roughly $25.5 million, including a $20 million Series A led by Leaders Fund. The company started with an enterprise platform for building and evaluating LLM applications, then extended into personal assistants and opened that codebase in April. Every assistant carries its own identity, memory, filesystem, and email address, and stays reachable through Slack, Telegram, email, and phone alongside the first-party apps. Model routing spans hosted models from Anthropic, OpenAI, and Google as well as local models through Ollama, with paid plans starting at 30 dollars a month. SPONSORED Download Vellum Mobile on IOS now [Learn more ](https://apps.apple.com/us/app/vellum-assistant/id6759934423?ref=testingcatalog.com) Putting voice, camera, scheduling, and remote machine control on the device that stays in a pocket moves the assistant into the hours of a day spent away from a desk. ### Z.ai launches GLM-5.3-Flash under MIT license URL: https://www.testingcatalog.com/z-ai-launches-glm-5-3-flash-under-mit-license/ Last updated: 2026-08-26T15:40:35.000Z Z.ai has launched GLM-5.3-Flash, the first natively multimodal model in its GLM-5 family, as a lower-cost option for coding, agent tasks, and visual reasoning. The mixture-of-experts model has 320 billion total parameters and 18 billion active parameters, down from 32 billion active parameters in GLM-4.5\. Z.ai says it outperforms GLM-5.2 across reported coding and agentic tests at one-tenth the price while approaching Claude Opus 4.8 on its internal coding benchmark. Trained on a 30-trillion-token multimodal corpus, GLM-5.3-Flash combines linear attention for local dependencies with sparse attention for relevant global context. At context lengths reaching one million tokens, IndexPool compresses groups of indexer key vectors to limit latency and memory use. Z.ai reports three times less attention compute and a 4.4-fold reduction in KV cache size compared with GLM-5.3. > Introducing GLM-5.3-Flash > > \- Leading capabilities at a highly competitive price > \- Natively multimodal with a 1M-token context window > \- A 320B-A18B model released under the MIT License > \- Previously previewed as Ox Alpha, running entirely on Chinese AI chips > > Blog:… [pic.twitter.com/KOCG4dkay3](https://t.co/KOCG4dkay3?ref=testingcatalog.com) > > — Z.ai (@Zai\_org) [August 26, 2026](https://x.com/Zai%5Forg/status/2092616204787626030?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) On Artificial Analysis Intelligence Index v4.1.1, the model scored 57 at a discounted cost of $0.045 per task, according to Z.ai. It also reached 63.4 on DeepSWE v1.1 against GLM-5.2's 46.2, and 48.8 on AutomationBench against 26.2\. Because the evaluations use different harnesses, context limits and generation settings, comparisons depend on each test setup. Visual reasoning is central to the release. Z.ai trained the model to inspect rendered interfaces, gameplay and 3D output, then assess and revise its work from visual feedback. The same approach covers documents, spreadsheets, presentations, dashboards and meeting materials, allowing the model to reason across text, images and structure. > GLM-5.3 Flash, aka Ox Alpha, has officially arrived. The announcement and benchmarks are up already, and the Pareto Frontier has been redrawn. [https://t.co/jVij8PJ5xc](https://t.co/jVij8PJ5xc?ref=testingcatalog.com) [pic.twitter.com/Z0Ug3nmnmL](https://t.co/Z0Ug3nmnmL?ref=testingcatalog.com) > > — Andrew Curran (@AndrewCurran\_) [August 26, 2026](https://x.com/AndrewCurran%5F/status/2092623826987491487?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Before launch, GLM-5.3-Flash appeared anonymously as ox-alpha on OpenCode and OpenRouter. Z.ai says it became the most popular model of the week on those services, with traffic served on Chinese AI chips. The company built an SGLang-based stack that separates encoding, prefill and decoding, reporting a threefold gain in end-to-end serving performance across tens of thousands of domestic accelerators. GLM-5.3-Flash is now available to all GLM Coding Plan users with three times the usable quota of GLM-5.3\. Its multimodal capabilities are offered in ZCode through Browser Use and Computer Use. The weights are available on [Hugging Face](https://huggingface.co/zai-org/GLM-5.3-Flash?ref=testingcatalog.com), and local deployment supports SGLang, vLLM and TokenSpeed. [Source](https://z.ai/blog/glm-5.3-flash?ref=testingcatalog.com) ### Claude unifies memory across Chat and Cowork URL: https://www.testingcatalog.com/claude-unifies-memory-across-chat-and-cowork/ Last updated: 2026-08-26T14:53:05.000Z Anthropic is rolling out a single Claude memory that follows users between chat and Claude Cowork, giving both products saved context when a task begins. The change starts today for people moving between conversations and Cowork’s cloud tasks. Memory is on by default for Free, Pro, and Max users on web, desktop, and mobile. Team and Enterprise access remains under administrator control, and individual users must turn it on. > Claude now has one memory across chat and Claude Cowork, and you decide what's in it. > > Hand Cowork a task and it starts from what Claude already knows from your chats: the project you talked through, your manager's preferences, or the client from last quarter. [pic.twitter.com/ViKapsT8K6](https://t.co/ViKapsT8K6?ref=testingcatalog.com) > > — Claude (@claudeai) [August 25, 2026](https://x.com/claudeai/status/2092299704864284888?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Cowork can now draw on details accumulated in chat, such as project priorities, team terminology, or preferences for manager updates. Information that arises while Cowork handles a task can carry back into later chats. A user could discuss an event agenda in chat, then ask Cowork to prepare its budget and logistics without repeating the city, headcount, or speakers. Claude now adds topics to memory as conversations happen instead of waiting until a conversation ends to produce a summary. A moved deadline can therefore appear in the next conversation without an explicit request to remember it. Users can pause or reset memory at any time. The stored information appears as short files organized by topic in Settings > Memory. Each file can be read, edited, or deleted, so correcting an outdated company name once changes what chats and Cowork tasks receive. Shared memory is inspectable, giving users control over carried context. Sensitive topics remain excluded by default, including health, race, ethnicity, religion, politics and gender identity. Users can opt in to saving them and will see a notice when Claude records one, but the choice applies only going forward. Claude will still refuse to store sensitive identification numbers, criminal history, immigration status, or material that violates Anthropic’s Acceptable Use Policy. The option can be turned off again at any time. On iOS and Android, users need the latest app version. Anthropic, the company behind [Claude](https://www.testingcatalog.com/tag/claude/), is using the release to connect its conversational product with Cowork’s cloud task runner through a common context layer. The result is continuity across where users work with Claude, paired with controls over what is stored and what remains excluded. [Source](https://claude.com/blog/claudes-memory-works-everywhere-and-you-decide-whats-in-it?ref=testingcatalog.com) ### Antigravity 2.0 adds new VCS controls and a terminal URL: https://www.testingcatalog.com/antigravity-2-0-adds-new-vcs-controls-and-a-terminal/ Last updated: 2026-08-25T15:39:45.000Z Google has added Git version-control controls and an embedded terminal inside Antigravity 2.0, putting repository review and command-line work in the coding assistant’s right-side panel. The update targets developers pairing with AI agents who need a reliable picture of every workspace change, not just edits made through an agent’s editing tool. The revised sidebar connects directly to the Git working tree and brings together three views: Agent Edits, Uncommitted, and Branch. It can show changes made by the agent, a separate editor, a Python program, or a Bash script, along with branch differences against origin/main. That closes a gap in the prior interface, where users often had to leave Antigravity to run git status or git diff before they could trust what the review pane showed. > No longer switch from Antigravity 2.0 to a separate terminal! Natively manage git workflows and use a native terminal for one-off bash commands. > > This has been a common ask that makes even more sense to ship now, given that models like Gemini 3.7 Flash have hit a level of… [pic.twitter.com/NxU5Ca4lau](https://t.co/NxU5Ca4lau?ref=testingcatalog.com) > > — Anshul Ramachandran (@\_anshulr) [August 25, 2026](https://x.com/%5Fanshulr/status/2092041143026622949?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Developers can open modified files, search within diffs, and switch between split and unified layouts. A staged-changes view supports selecting files, staging or unstaging them, drafting a commit message, and committing from the sidebar. Antigravity can also preview an automatically generated message before creating the commit. The adjacent Terminal tab handles the rest of the workflow without leaving the app. Users can run tests, linters, builds, package commands, and other utilities, with Google’s demo highlighting git status and go test ./... commands. Together, the terminal and VCS pane let teams inspect agent output, verify side effects from other tools, and move changes toward a commit from one workspace. The [Antigravity](https://www.testingcatalog.com/tag/antigravity/) Team is framing the release around reviewability and trust in AI-assisted development. Antigravity 2.0 previously surfaced agent edits in its sidebar, while the new Git-backed view reflects the wider state of the branch and working directory. The accompanying interactive simulation lets users test view switching, diff inspection, terminal commands, and commit flow, although Google notes that its appearance may differ slightly from the product. [Source](https://antigravity.google/blog/vcs-and-terminal?ref=testingcatalog.com) ### ICYMI: SpaceXAI plans NVIDIA Vera deployment for Starmind URL: https://www.testingcatalog.com/icymi-spacexai-plans-nvidia-vera-deployment-for-starmind/ Last updated: 2026-08-25T15:35:56.000Z NVIDIA says SpaceXAI will deploy its Vera CPUs to power the next generation of agentic AI workloads behind Grok, linking the chipmaker’s new CPU architecture to SpaceXAI’s push toward gigawatt-scale computing. The plan also reaches beyond terrestrial data centers: SpaceXAI intends to carry an optimized Vera Rubin NVL72 foundation into orbit through its first-generation Starmind AI satellite. Vera is designed for the CPU-heavy work that surrounds model inference, including tool orchestration, code execution, data processing and simulation. Those tasks take place between model calls, so faster CPU processing can help agents act sooner while keeping attached GPUs occupied. NVIDIA says the processor uses 88 company-designed Olympus cores, Spatial Multithreading and high-bandwidth LPDDR5X memory capable of up to 1.2TB/s. > SpaceX, in partnership with Nvidia, has designed a space-optimized Vera Rubin NVL72 system for launch to orbit in Q4 next year, with significant scale in 2028 [https://t.co/qdDq8YBkzl](https://t.co/qdDq8YBkzl?ref=testingcatalog.com) > > — Elon Musk (@elonmusk) [August 24, 2026](https://x.com/elonmusk/status/2091939113008238838?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) NVIDIA claims Vera can complete tasks up to 1.8 times faster than x86 CPUs across agentic AI, reinforcement learning and data-processing workloads. The company presents it as the first CPU purpose-built for AI agents, though the claim comes from NVIDIA. SpaceXAI president Mike Nicolls said the performance and memory bandwidth are intended to handle large volumes of orchestration, code, and data work while leaving GPUs focused on their primary jobs. At the system level, SpaceXAI plans to expand the infrastructure supporting Grok around the Vera Rubin platform. The architecture combines accelerated computing, NVLink, Spectrum-X Ethernet, BlueField data processing and NVIDIA software across training, reasoning and inference. NVIDIA’s pitch is that a common stack can serve both large AI factories on Earth and systems adapted for orbital conditions, where power, heat, bandwidth, reliability and physical integration create different limits. The proposed Starmind satellite would use an optimized Vera Rubin NVL72 rack-scale system, with NVIDIA and SpaceXAI adapting it for orbital computing while retaining the same software ecosystem. That would give SpaceXAI one computing foundation spanning agent workloads, Grok infrastructure, and planned space-based AI. NVIDIA did not announce a deployment date, and it cautioned that the products and features remain in development, will be offered only if and when available, and could change before release. [Source](https://nvidianews.nvidia.com/news/spacexai-adopts-nvidia-vera-cpu-to-accelerate-agentic-ai-at-massive-scale?ref=testingcatalog.com) ### Grok Bot to get Templates and multi-account support soon URL: https://www.testingcatalog.com/grok-bot-to-get-templates-and-multi-account-support-soon/ Last updated: 2026-08-24T19:54:28.000Z TestingCatalog found that xAI is already preparing several new features for the recently released Grok Bot, including templates, multi-account support, Chrome profile importing, and new networking controls. One major addition is a new Templates tab in Settings, where users can browse shared templates! > 🤖SpaceXAI is working on Shareable Templates for Grok Bot. > > Ask a Grok Bot to pack itself. It stages a private draft, you publish a share link, and the other person imports simply by going to the link and clicking Add to Grok Bot > > This is so cool! We can't wait till this ships… [https://t.co/nrfMxgu8vc](https://t.co/nrfMxgu8vc?ref=testingcatalog.com) [pic.twitter.com/3Y2OHV712D](https://t.co/3Y2OHV712D?ref=testingcatalog.com) > > — ️ (@blankspeaker) [August 24, 2026](https://x.com/blankspeaker/status/2091969780706337090?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Instead of building a Bot from scratch, users could start with a predefined role or configuration. That would make templates broader than reusable task automation and could provide ready-made setups for roles such as researchers, sales assistants, or operations agents. ![Grok](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-24-at-01.12.25.webp) Multi-account support is also in development, allowing users to add several accounts to the Grok Bot app and switch between them. This could be particularly useful for separating personal, work, or organization accounts. ![Grok](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-24-at-01.11.33.webp) Several additional settings are currently hidden. Chrome Profile Import would let users transfer Chrome profiles with cookies, potentially letting Grok Bot access websites using existing authenticated browser sessions. Desktop Egress Routing would route the Bot’s network traffic through the desktop running Grok Bot via an egress tunnel. xAI’s documentation already describes a beta setting for routing cloud-computer traffic through a user’s computer, including as a workaround for websites that block datacenter IP addresses. ![Grok](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-24-at-00.31.30.webp) Grok Bot launched in beta on August 11 as xAI’s persistent agent product, with Bots operating through a cloud computer and working across apps and websites. xAI expanded access on August 21 to SuperGrok Plus, Cursor Pro+, and all Cursor Teams plans. SuperGrok Plus currently costs $100 per month. > We're making Grok Bot more widely available. > > All SuperGrok Plus, Cursor Pro+, and Cursor Teams subscribers now have access. > > We're also offering a free trial with limited usage for all other users. [pic.twitter.com/0D3oUDZQCM](https://t.co/0D3oUDZQCM?ref=testingcatalog.com) > > — Grok Bot (@bot) [August 21, 2026](https://x.com/bot/status/2090852881373311369?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) These upcoming controls suggest xAI is now building infrastructure around reusable Bot configurations, account management, authentication, and more flexible network access. ### Google prepares Gemini App for Avatars, Plugins, and Gemini 4 URL: https://www.testingcatalog.com/google-prepares-gemini-app-for-avatars-plugins-and-gemini-4/ Last updated: 2026-08-24T14:01:27.000Z Google continues building out its Gemini desktop app, and the latest version includes several additions that point to broader expansion. Avatars are being added to the desktop app, with a dedicated Settings section where users can create and manage them directly from the app. Google is also bringing over the previously spotted “[Customize](https://www.testingcatalog.com/google-is-working-on-plugins-for-gemini-enterprise/)” area. It remains hidden, as on the web, but opens a discovery screen for apps, skills, and plugins. The direction resembles what rivals such as Claude and Codex are already pursuing with their desktop apps. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-24-at-00.43.59.webp) Google is relatively late to this approach. Broad connector and MCP support is still missing from the consumer Gemini app, potentially because Gemini models may require additional model-level work before they can reliably use a wider range of external tools. Connectors already exist in Gemini Business with a limited selection, while Spark was the first Google product to include custom MCP support. That capability still has not reached the main consumer experience. At the same time, Google is preparing internally for Gemini 4\. Based on earlier information about roughly six-month training cycles, the next major model could arrive around December. The summer cycle was originally expected to produce Gemini 3.5 Pro, but with that release now canceled, Google may move directly to Gemini 4. > Google is finally gearing up to release Gemini 4\. > > \~2-3 days ago, \[REDACTED\] had 120+ new mentions of Gemini 4, up from 0\. Less than 24 hours ago, \[REDACTED\] product got 6 new mentions similar to the other mentions. > > It is starting to propagate. > > — Bedros Pamboukian (@bedros\_p) [August 23, 2026](https://x.com/bedros%5Fp/status/2091666760059793534?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) That timing would resemble [Gemini 3](https://www.testingcatalog.com/gemini-3-references-found-in-code-hint-at-google-deepminds-next-ai-model/), whose early checkpoints began appearing during the summer before the model arrived later in the year. The desktop app is already being adapted for new types of responses, including Gmail and Calendar widgets, suggesting Google is preparing the client to handle upcoming Gemini 4 capabilities. For now, these desktop changes remain in development, and the new model-related behavior does not appear to be in testing outside Google. The eventual combination of avatars, apps, skills, plugins, connectors, and richer response widgets could turn [Gemini](https://www.testingcatalog.com/tag/gemini/) desktop into a much broader AI workspace, but the rollout timeline remains unclear. ### Exclusive: Early outputs of Muse Video model from Meta URL: https://www.testingcatalog.com/exclusive-early-outputs-of-muse-video-model-from-meta/ Last updated: 2026-08-19T23:04:33.000Z Meta is currently testing its upcoming Muse Video model in beta with selected partners. The model, labeled “Beta,” was [first previewed in July](https://www.testingcatalog.com/meta-launches-muse-image-across-its-apps-and-previews-muse-video/) alongside Muse Image. At that time, Meta shared initial samples and announced that Muse Video would soon be available to creators and Meta AI. TestingCatalog recently gained access to Muse Video and tested it directly. Based on our early generations, the model shows state-of-the-art potential, particularly in fine detail, world understanding, and temporal consistency. It currently produces 10-second videos, but the output quality is remarkably high. Meta has also confirmed that Muse Video supports native audio, although the company previously acknowledged remaining gaps around audio-video synchronization and physically accurate fast motion. Muse Video samples 👀 0:00 /0:10 1× "Cyberpunk hacker robot working in front of many monitors" 0:00 /0:10 1× "A shaky, low-quality iPhone video recording. A guy holding the camera selfie-style walks up to random people on a busy city street and asks each one, "How do you feel about being AI generated?" The video has authentic smartphone artifacts—slight motion blur, overexposed sky, wind noise, and the occasional finger over the lens." 0:00 /0:10 1× "The rocket at the rocket station is launching. Everything looks normal until the rocket explodes into confetti pieces and leaves nothing behind." 0:00 /0:10 1× "The man couldn't believe it when he finally became a kitten dad. Excitedly he shouted "I'm a dad now! A cat dad!". He's so happy that he starts tearing up. The kitten meows before the video ends." 0:00 /0:10 1× "A cat made out of Lego technic in the rain made out of fluffy plastic rain" 0:00 /0:10 1× "A glass of red wine tips over on a white marble counter, spilling toward the camera. Static shot" 0:00 /0:10 1× "A woman with curly red hair and a green scarf walks toward camera down a crowded Tokyo street at night, stops, and looks up." 0:00 /0:10 1× "Orbiting shot around a chef plating a dish, 180-degree arc, shallow depth of field." 0:00 /0:10 1× "A neon sign reading ‘OPEN 24 HOURS’ flickers on above a diner door in the rain." 0:00 /0:10 1× "A car’s headlights sweep across a bedroom wall at night, shadows shifting." 0:00 /0:10 1× Two swordsmen in a bamboo forest exchange three strikes, blades meeting mid-frame. Wide shot, overcast light. 0:00 /0:10 1× "Two fighters in a dimly lit warehouse; one throws a punch, the other blocks and counters. Handheld camera, shallow depth of field." There is still no confirmed public release date. Meta has already stated that Muse Video is coming to Meta AI, making Meta.ai and the Meta AI app the most obvious initial destinations. Given that Muse Image is available for free everyday creation, Muse Video could follow a similar model, although its pricing or generation limits remain unknown. The broader distribution possibilities are particularly important. Muse Video could become part of Meta’s Vibes feed for AI-generated videos and later expand across Instagram, Facebook, and other Meta products. Instagram’s video-heavy ecosystem is an obvious fit, while Meta’s Edits video creation app could also benefit from native generative video capabilities. For Meta, video generation fits directly into its wider strategy of building its own media models and distributing them across products used by creators and consumers. Facebook and Instagram already revolve heavily around visual content, giving Meta a natural distribution advantage if Muse Video proves competitive at scale. Our current access indicates that development has moved beyond Meta’s original public preview into working beta testing. Don't forget to take a look at the video samples we have collected so far. Test prompts from [@MarsForTech](https://x.com/MarsForTech?ref=testingcatalog.com), [@bughuntergeek](https://x.com/bughuntergeek?ref=testingcatalog.com), [@BartokGabi17](https://x.com/BartokGabi17?ref=testingcatalog.com) ### Meta launches AI desktop app for macOS with screen sharing URL: https://www.testingcatalog.com/meta-launches-ai-desktop-app-for-macos-with-screen-sharing/ Last updated: 2026-08-19T21:24:32.000Z As we [reported earlier](https://www.testingcatalog.com/meta-prepares-desktop-ai-app-for-macos-as-rivals-race-ahead/), Meta has been preparing a dedicated Meta AI desktop app for macOS. It's now available in [certain regions](https://ai.meta.com/meta-ai/assistant/?ref=testingcatalog.com), and we finally got a chance to test it. Meta officially announced the Mac app on August 19, confirming screen sharing and system-wide dictation as core desktop features. The app looks very similar to Meta AI on the web and offers largely the same functionality, but desktop integration adds several notable capabilities. Users can attach a specific screen or window directly from the prompt bar and use its contents as context. Voice dictation is also available, including through a system-wide shortcut that can summon a compact Meta AI prompt bar from anywhere on macOS. 0:00 /0:16 1× The dictation feature is particularly interesting from a model perspective. It does not necessarily indicate that Meta is preparing an entirely new speech-to-text model, as the company already has public speech-recognition technology, including Omnilingual ASR. Still, bringing dictation directly into Meta AI could provide a consumer-facing surface for newer speech technology in the future. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-19-at-21.55.22.webp) The build we accessed is already labeled version 1.0, suggesting the desktop product has moved well beyond an early prototype. For Meta, bringing Meta AI to macOS is another step toward closing the desktop gap with OpenAI, Anthropic, Google, and other competitors. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-19-at-21.55.27.webp) There is still considerable room to catch up. Artifacts are already supported, but Meta AI itself lacks the deeply integrated coding workflows available through ChatGPT and Claude. Meta separately launched Muse Code and Muse Spark 1.2 earlier this month, so the missing piece is increasingly integration rather than underlying coding capability. Browser use and computer control remain other major gaps in the consumer desktop app, even as Meta experiments with computer-use capabilities through its developer stack. ### Fabraix opens public access for Nyx AI security red team agent URL: https://www.testingcatalog.com/fabraix-opens-public-access-for-nyx-ai-security-red-team-agent/ Last updated: 2026-08-19T18:48:35.000Z Fabraix has opened a second public launch of Nyx, its autonomous red-teaming agent, positioning the product as continuous security testing for teams shipping customer-facing AI. Nyx takes an endpoint, a URL, or a phone number as a target, configures its own test plan from a description, and runs adversarial attacks against chat, voice, browser, and coding agents on a repeating schedule. ![Attack samples](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/image-3.png) The agent works from a library of more than 10,000 jailbreaks and attack strategies collected from public records, using them as starting points for multi-turn attacks that adapt in real time to how a target defends itself. Testing runs against the agent directly and indirectly through its environment, with payloads placed inside webpages, documents, files, messages, and tool outputs hosted on controlled replicas of SaaS products that Fabraix maintains. Nyx operates as a pure blackbox and requires no source code, model weights, credentials, or network access, with each run kept isolated and ephemeral and never trained on. Every finding arrives with the attack steps, the agent responses, and the failure that resulted, so the same test can be replayed after a release to check whether the fix held. Fabraix reports that Nyx often surfaces a first vulnerability within minutes or hours, and cites a 78 percent attack success rate on AgentHarm, a benchmark for offensive AI security. Access runs through a CLI installed via npm alongside a REST API, with each audit defined in a YAML file covering target, objective, and budget, and results streaming live during a run. Token-based authentication covers CI pipelines, so audits can sit inside an existing release process instead of arriving as a one-off engagement. Fabraix also maintains a public Playground where the community attempts to break live agents with published system prompts, and publishes research alongside the product, including Adversarial Cost to Exploit, a benchmark that scores AI security by the token spend an attacker needs to breach an agent. SPONSORED Demo booking is available now [Learn more ](https://fabraix.com/?ref=testingcatalog.com) Fabraix was founded in 2026 by Ahmed Aly and Ibrahim Abdu and is backed by Y Combinator in its Summer 2026 batch. Abdu built AI agents at Meta that diagnosed and fixed production errors, after earlier work on compilers and database engines in fintech. Aly led the international payments fraud team at Monzo and was the first data scientist at a Sequoia-backed startup, where he built a fraud engine that grew to process more than a billion dollars in annual B2B transactions. The company reports that Nyx has already found exploitable failures in public-facing agents run by dozens of Fortune 500 companies. The intended audience is security leaders and internal red teams at mid-size and enterprise companies putting agents in front of customers, along with the engineers currently doing that testing by hand. Access is arranged through a demo request, and no public pricing has been listed. ### Ornith-1.5 open models launch in 397B, 35B, and 9 B sizes. URL: https://www.testingcatalog.com/ornith-1-5-open-models-launch-in-397b-35b-and-9-b-sizes/ Last updated: 2026-08-19T14:30:50.000Z Ornith has released Ornith-1.5, a family of open models that extends the self-scaffolding framework from Ornith-1.0 into a closed self-improvement loop. Where the previous generation wrote the scaffold around a fixed set of human-curated tasks, Ornith-1.5 proposes the tasks themselves, generates a task-specific scaffold for each one, and produces the solution rollouts used for reinforcement learning. The release covers three scales: a 397B mixture-of-experts flagship, a 35B mixture-of-experts model activating 3B parameters per token, and a 9B dense model shipping with a quantized Mobile build for iPhone and Android. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/image-1.png) Each training cycle runs in three stages. Given an environment or codebase, high-level instructions about the task type, and the model's own history of solved problems, the system proposes progressively harder tasks that sit beyond what it has already handled. It then generates or refines a scaffold covering instructions, tools, decomposition strategy, and orchestration, and produces a rollout conditioned on both. Reward propagates back across all three stages, so the system learns to write better solutions, more useful training tasks, and more reliable evaluation harnesses at once. Task reward multiplies three signals: whether the task and scaffold form a valid and verifiable environment, whether difficulty sits near the current capability frontier, and whether the task is novel against work already generated. Frontier difficulty targets a 0.2 empirical success rate, so a task loses value to the generator once the model starts clearing it reliably. Validity acts as a hard gate that zeroes out malformed tasks, and all three stages are optimised with GRPO. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/image-2.png) On the company's published tables, averaged over five independent runs, Ornith-1.5-397B scores 85.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, which Ornith reports as on par with Claude Opus 4.8 at 85.0 and 59.0 and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731 at comparable scale. The 35B reaches 68.5 and 79.0 on Terminal-Bench 2.1 and SWE-Bench Verified while activating 3B parameters per token, and the 9B reaches 47.0 and 70.6, which the company places above Gemma 4-31B and Qwen 3.6-35B. Coverage extends past coding into reasoning and agentic work, with 92.8 on GPQA Diamond and 86.6 on BrowseComp at flagship scale. SPONSORED Check out the weights and full tables [Learn more ](https://ornith.ai/ornith%5F1%5F5.html?ref=testingcatalog.com) Ornith is the model line from DeepReinforce, the research team that shipped Ornith-1.0 in June 2026 across 9B dense, 31B dense, 35B MoE, and 397B MoE variants under an MIT license, with weights on Hugging Face. That release was post-trained on Gemma 4 and Qwen 3.5 checkpoints and introduced the idea of treating the scaffold as a learnable object co-evolving with the policy. DeepReinforce has published reinforcement learning optimisation research in the open before, including CUDA-L1 and the IterX agent loop, and reward hacking has been a running concern across that work. Ornith-1.5 carries that defence into task generation, where an unverifiable task earns nothing. HuggingFace: [https://huggingface.co/collections/ornith-ai/ornith-15](https://huggingface.co/collections/ornith-ai/ornith-15?ref=testingcatalog.com) ### Google tests Computer Use on Gemini Desktop URL: https://www.testingcatalog.com/google-tests-computer-use-on-gemini-desktop/ Last updated: 2026-08-18T15:29:55.000Z Google continues preparing its Gemini desktop app to close the gap with competitors. The company has finally started working on computer-use functionality, which would allow Gemini to control other apps, access files in selected folders, and perform tasks similar to those currently supported by Claude and ChatGPT. Traces of the feature have appeared in the app’s settings, suggesting Google may have begun testing it with a limited group of trusted testers. Computer use will likely arrive alongside a slightly updated Gemini Spark interface. Instead of being activated through a toggle in the navigation bar, it is expected to appear as a “Task” option in the prompt-bar selector, similar to the existing controls for generating images and videos. While Gemini is operating the computer, users will be able to follow its actions through a mini-view of the active screen. The app is also expected to include practical controls, such as preventing the device from going to sleep while a task is running. Another notable addition is an optional safety feature that automatically backs up selected folders to Google Drive before Gemini begins a task. This could provide a useful safety net for people who do not use Git or another version-control system. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-18-at-00.53.24.webp) However, the feature is also a double-edged sword. Its existence suggests Gemini could potentially make unwanted changes to a workspace, while the backup process raises practical questions. If a selected folder contains hundreds of gigabytes, uploading it to Google Drive could take a long time and consume substantial storage before the task can even begin. Computer use appears to be Google’s next major frontier for [Gemini](https://www.testingcatalog.com/tag/gemini/). Its value will depend not only on what the agent can do, but also on how Google handles permissions, recovery, large folders, and user oversight. ### Anthropic tests Hub Mode in Claude for sub-agents URL: https://www.testingcatalog.com/anthropic-tests-hub-mode-in-claude-for-sub-agents/ Last updated: 2026-08-26T18:52:54.000Z Anthropic is working on a new feature called Hub Mode, recently spotted in its mobile app. The option is marked as internal, suggesting it remains under internal testing and has not been released more broadly. Hub Mode appears next to a planet icon. A similar icon [previously appeared under the name Orbit](https://www.testingcatalog.com/anthropic-is-working-on-orbit-its-upcoming-proactive-assistant/) before being removed, raising the possibility that Hub Mode is a new version of the same concept. The current description says the mode will apply to new tasks and requires users to restart the app after activation. However, enabling it did not produce any visible changes on our side, suggesting that the underlying functionality is not yet operational. > Anthropic is working on a Hub Mode 👀 > > \* Speculation - It might be related to an interface where Claude will be operating multiple sub-agents and visualizing the overall progress to the user. [pic.twitter.com/N4aIgmXAHD](https://t.co/N4aIgmXAHD?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 18, 2026](https://x.com/testingcatalog/status/2089613179936555491?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The only related reference on the web appears within Claude Code. Opening the Hub destination leads to an interface containing a list of sub-agents. Selecting one opens a page stating that [Claude](https://www.testingcatalog.com/tag/claude/) will generate a task board for the user and their agents. No other functional details are currently visible. This could point to a new interface for managing agents or sub-agents, potentially resembling the recently released Grok Bot, where users can create multiple specialized agents. However, there is no confirmed connection between Hub Mode in the Claude mobile app and the Hub interface found in Claude Code on the web. For Anthropic, such a feature would fit its broader push toward agent-based workflows across Claude Code, Cowork, and collaborative projects. If the two Hub references are connected, users could eventually receive a central place to launch tasks, monitor specialized agents, and coordinate their work. For now, its exact purpose, limitations, target users, and release timeline remain unclear. ### Cursor begins Origin code hosting rollout for paid plans URL: https://www.testingcatalog.com/cursor-begins-origin-code-hosting-rollout-for-paid-plans/ Last updated: 2026-08-18T14:57:10.000Z Cursor has begun rolling out Origin, a code-hosting service built directly into [Cursor](https://www.testingcatalog.com/tag/cursor/). The early beta began rolling out to users on all paid plans on August 17, 2026, except enterprise organizations whose administrators opt out. Origin launches with repositories, pull requests, code browsing, GitHub sync, and access to Cursor agents, giving users one place to store code, review changes, and ask agents to work across a project. > Origin, our code hosting platform, is now live. > > It's fast, easy to use, and deeply integrated with Cursor. > > Get started by syncing your repos from GitHub. [pic.twitter.com/aqRHavAOQg](https://t.co/aqRHavAOQg?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [August 17, 2026](https://x.com/cursor%5Fai/status/2089399057659596847?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The new Codebase tab is the entry point for Origin repositories. Users select +New, name a repository, and receive CLI installation instructions plus commands for cloning it or pushing an existing local project. Once pushed, Cursor hosts the project. Each codebase name also becomes part of every repository URL. Teams do not have to move projects away from GitHub. They can connect an organization, choose repositories to sync, and place those projects beside Cursor-hosted repositories. Synced copies update in real time and can be browsed, searched, and pulled from Origin, while GitHub remains the source of truth and continues to receive pushes. Anyone with read or write access to a synced repository can view it in Cursor, and you can disconnect the sync at any time. Pull request pages bring together the timeline, commits, checks, changed files, diff, comments, and merge controls. On synced projects, comments written in Cursor appear on GitHub, while GitHub reactions and replies show up in Cursor within seconds. Reviews assigned through GitHub can also be completed and merged from Cursor. Agents sit alongside the code and pull requests. Users can ask questions about the code they are viewing, request changes, update a pull request, or have Cursor push a branch. Repository settings expose sync status, access controls, and connected apps. ![Cursor](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Cursor-Agent-Turn-your-ideas-into-code-08-15-2026_12_10_AM.jpg) Cursor is also building an app ecosystem around Origin. Vercel can create a preview deployment for every pull request and ship the change after a merge. Depot and Buildkite can run existing GitHub Actions workflows, while Buildkite also supports native pipelines. Cursor says these are the first parts of a system designed for agent scale, with more agent-native features and app connections planned. [Source](https://cursor.com/changelog/origin-code-hosting?ref=testingcatalog.com) ### Vorflux opens sign-ups for its automated coding platform URL: https://www.testingcatalog.com/vorflux-opens-sign-ups-for-its-automated-coding-platform/ Last updated: 2026-08-17T15:28:44.000Z Vorflux has opened public sign-ups for a cloud platform that takes a software task from intent to a merged pull request without a laptop staying in the loop. The company positions the product as an autopilot for software engineering: a session starts from what Vorflux describes as two links and a sentence, then covers planning, building, testing, review, and the merge itself, with the person who filed the task approving by exception. Sign-up is live on Vorflux's website, and the Vorflux team configures larger setups in collaboration with the customer. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/image.png) The machine is what separates it from a cloud sandbox. Each session spins up an entire stack on a dedicated EC2 instance, cloning repos, installing libraries, and bringing up frontend, backend, mobile services, and databases, with CI/CD, feature flags, and environment variables wired in. Vorflux then snapshots the box with everything live, so later sessions wake on a computer that is already running. Because the application genuinely runs, a testing pass can drive a browser through the real user flow, record it, and attach that recording to the pull request as proof. A phone boots in the cloud for native tests, and every branch gets a live URL for review before merge. > Circle the bug. Send an-inline request to Vorflux. Get back a PR with proof it's fixed. [pic.twitter.com/Pf3Ru1qyY6](https://t.co/Pf3Ru1qyY6?ref=testingcatalog.com) > > — Vorflux (@vorfluxai) [August 11, 2026](https://x.com/vorfluxai/status/2087268055521149407?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Vorflux trains none of its own models and routes each stage to whichever frontier model leads that job, covering Claude, GPT, Gemini, and open-weight options. Planning receives the largest token budget, expanding intent into a plan of sub-tasks, each with its own test cases, with a model from a separate lab critiquing the draft before code is written. The same cross-lab logic runs at review, where a reviewer outside the author model's family grades the diff and retells it in plain language with the most important change first. A merge queue then handles rebases and conflicts against a moving master. > Ask why prod is slow. Get back a diagram of exactly where [pic.twitter.com/qoB5mj6ETn](https://t.co/qoB5mj6ETn?ref=testingcatalog.com) > > — Vorflux (@vorfluxai) [August 6, 2026](https://x.com/vorfluxai/status/2085499154172903436?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Billing splits three ways: 1. Customer provider keys paid directly. 2. Tokens bought through Vorflux at standard provider rates with no platform fee. 3. An existing Codex subscription routed in under an OpenAI partnership. Compute runs hourly, from roughly 21 cents for a 2 vCPU box up to 3.39 dollars for 32 vCPUs, and servers auto-stop after an hour of idle time. A white-glove tier embeds a Vorflux engineer who runs the agents against the customer backlog and ships verified pull requests weekly. SPONSORED Start testing Vorflux [Learn more ](https://devtoolsacademy.link/tc-s?ref=testingcatalog.com) Vorflux was founded by Prasanna Sankar, co-founder and former chief technology officer of Rippling, and launched in July 2026 with funding from Y Combinator, Peak XV Partners, Powerset, and Alliance, alongside angels including Parker Conrad, Immad Akhund, and Balaji Srinivasan. Sankar has framed human time, not tokens, as the remaining bottleneck in software delivery, and the company builds around a layer it calls the harness: the place where a team encodes its own engineering standards so they run on every session. The stated target is technical founders, CTOs, and engineering leaders carrying review backlogs and testing debt. ### Anthropic develops Slack-like features for Claude Desktop URL: https://www.testingcatalog.com/anthropic-develops-its-own-slack-as-claude-code-projects/ Last updated: 2026-08-16T14:42:27.000Z Anthropic has updated its upcoming [collaborative projects for Claude Code](https://www.testingcatalog.com/anthropic-develops-claude-driven-managed-projects/). During project creation, users can now add repositories directly as persistent context. The interface recommends adding only the repositories required for each session, and suggests attaching additional repos on a session-by-session basis when needed. > ANTHROPIC 🔥: Claude Tag is coming to Claude Desktop in the form of collaborative Projects for Claude Code. > > That's essentially a built-in Slack 👀 > > \> Earlier spotted Managed Projects feature got a slight upgrade, mentioning that users will be able to spawn different threads… [pic.twitter.com/aOzCEK2m9Y](https://t.co/aOzCEK2m9Y?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 16, 2026](https://x.com/testingcatalog/status/2088999252559024420?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) At its core, the feature is designed around shared memory and context that persist across sessions, with multiple people able to work inside the same session. The latest description also introduces multiple threads within a session. A useful parallel is Slack: the project session effectively becomes a channel where several people and Claude are present, while threads allow separate sub-conversations around individual tasks. This fits Anthropic’s existing direction with [Claude Tag on Slack](https://www.testingcatalog.com/anthropic-launches-claude-tag-on-team-and-enterprise-plans/), where Claude can use channel and thread context alongside connected codebases. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Claude-Code-08-16-2026_04_04_PM--1-.jpg) Anthropic has suggested this collaborative model could represent the future of software development. Instead of traditional pair programming, or a developer working alone with AI, multiple developers and Claude can work together on the same project. Anthropic has reportedly used this approach internally for some time, and managed projects could bring that workflow directly into Claude without requiring Slack. > Great research went into making Claude proactive. Our data is showing unprompted messages are down \~45% -- Claude Tag is both more context-aware across your work and more attuned on when to keep quiet. > > Best part - you don't pay for Monitoring either![https://t.co/Tvo3347fY9](https://t.co/Tvo3347fY9?ref=testingcatalog.com) > > — Noah Zweben (@noahzweben) [August 13, 2026](https://x.com/noahzweben/status/2087985717817274698?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Anthropic's focus has been around Claude Tag recently Another important piece we discovered earlier is that [Claude](https://www.testingcatalog.com/tag/claude/) is expected to [maintain these projects itself](https://www.testingcatalog.com/anthropic-develops-claude-driven-managed-projects/). It can periodically revisit context files and memories to reorganize and refine them, turning the project into a persistent workspace rather than static shared chat history. Availability also appears to be expanding beyond friends-and-family or alpha testing, with early-access organizations potentially already having access. There is no confirmed public rollout date yet. This may look like a relatively small addition at first, but it could become one of Anthropic’s larger workflow releases. The company is increasingly positioning Claude as infrastructure for shared organizational work, rather than only a one-to-one coding assistant, and collaborative projects appear to be a direct extension of that strategy. ### Claude turns into Super App, adding web browser for Cowork URL: https://www.testingcatalog.com/claude-turns-into-super-app-adding-web-browser-for-cowork/ Last updated: 2026-08-16T14:15:26.000Z Anthropic continues to enhance Claude Desktop, with recent app developments suggesting two potential additions that could move the product toward becoming a comprehensive AI super app. TestingCatalog has identified work on a built-in web browser for Cowork sessions. While a similar feature exists within Claude Code, it is currently not easily accessible from Cowork. The forthcoming implementation aims to integrate browsing directly into Cowork sessions, allowing users to open web pages alongside an active task while Claude maintains the surrounding context. This enhancement would help Anthropic bridge a feature gap with OpenAI’s Codex, where browsing, coding, and agent workflows are increasingly integrated within the same desktop environment. Anthropic seems to view this as a broader opportunity for Claude Desktop, not merely an isolated browser feature. This move aligns with a broader trend towards AI desktop super apps. Microsoft is progressing in this direction, Google is developing a Gemini desktop app, Meta is working on its own desktop application, and xAI has already launched its Grok app. With most major AI companies advancing towards persistent desktop environments that combine agents, tools, and browsing, Anthropic faces clear competitive pressure to match these developments. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-08-16-at-15.32.45.png) Simultaneously, TestingCatalog has observed work on a dedicated comparison mode. This feature would allow users to manually switch between regular chats and a view designed to compare responses from different [Claude](https://www.testingcatalog.com/tag/claude/) models. For instance, users could compare Opus with Fable 5 to determine if the output difference justifies spending more credits. This does not appear to be an internal A/B testing interface. Instead, the comparison mode is being developed as a user-facing option that can be activated when needed and deactivated for normal conversations. Both features remain unreleased, and no launch date has been announced. They were identified through references appearing in recent Claude builds. ### Google launches Sheets canvas for Gemini mini-apps URL: https://www.testingcatalog.com/google-launches-sheets-canvas-for-gemini-mini-apps/ Last updated: 2026-08-13T21:22:35.000Z Google is rolling out Sheets canvas, a Gemini-powered feature that turns spreadsheet data into custom, interactive “mini-apps” inside Google Sheets. The new read-write layer sits directly on top of a spreadsheet, giving users a visual way to organize, edit and navigate information without formulas, programming, or a separate app. Users can open a spreadsheet, select “Create canvas” in the Ask Gemini side panel, and describe the result they want in natural language. Gemini builds a layout from the existing data, and follow-up prompts can alter its design, arrangement, or functionality. Edits made in either the canvas or the underlying sheet sync in real time. Because the canvas lives as a tab in Google Sheets, it can also be shared with collaborators like a regular sheet. > Staring at hundreds of rows and columns, wishing your spreadsheet was an interactive dashboard? 📊✨ > > Introducing Sheets canvas in Google Sheets—turning static data into interactive "mini-apps" with a simple prompt. Available now globally in English to Google AI Pro and Ultra… [pic.twitter.com/4GIHwL4Vm0](https://t.co/4GIHwL4Vm0?ref=testingcatalog.com) > > — Google Docs (@googledocs) [August 13, 2026](https://x.com/googledocs/status/2087964937746358463?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Google is positioning Sheets canvas for everyday projects where long rows and columns make key details harder to see. Students can turn assignment lists into visual study trackers with progress, reminders, and a view of the week ahead. Fantasy football players can create roster command centers for comparing player statistics and standings. Wedding planners can convert RSVP records into seating charts, move guests between tables, and keep the central spreadsheet current. The launch gives Gemini a new role in how people present and work with data already held in Sheets. Instead of requiring code or a separate tool, Sheets canvas lets people build a task-specific view through prompts while keeping the spreadsheet underneath it. Its two-way syncing is central to the approach: the visual layer is not a static presentation, and updates remain connected to the source data. Sheets canvas is now available globally in English to Google AI Pro and Ultra subscribers. Google has also begun rolling it out to Workspace customers on Business or Enterprise Standard and Plus plans, as well as subscribers to the Google AI Pro for Education add-on. Access begins inside an existing Google Sheets spreadsheet through the Ask Gemini side panel. [Source](https://blog.google/products-and-platforms/products/workspace/sheets-canvas-for-google-sheets-spreadsheets/?ref=testingcatalog.com) ### Google launches Gemini 3.7 Flash for coding and AI agents URL: https://www.testingcatalog.com/google-launches-gemini-3-7-flash-for-coding-and-ai-agents/ Last updated: 2026-08-13T21:18:35.000Z Google has launched Gemini 3.7 Flash, calling it its most intelligent workhorse model for coding and agents. Arriving three weeks after Gemini 3.6 Flash, the model targets developers, enterprises and subscribers who need production workflows with less oversight. Google says developer feedback and algorithmic advances shaped the release. > Introducing Gemini 3.7 Flash : ) > > \- it is fast! > \- 50% lower price than 3.6 flash (through end of year) > \- strong intelligence increase in only \~3 weeks (thanks to some awesome algorithmic improvements) > \- available in the API, AI Studio, Antigravity, and more! [pic.twitter.com/hxnd83CiXD](https://t.co/hxnd83CiXD?ref=testingcatalog.com) > > — Logan Kilpatrick (@OfficialLoganK) [August 13, 2026](https://x.com/OfficialLoganK/status/2087948481721962669?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) It posts gains in software engineering, web development and knowledge-heavy work. Gemini 3.7 Flash scored 43.6% on FrontierCode 1.1 Main, up from 34.4% for 3.6 Flash, and 65.3% on DeepSWE v1.1 versus 49.0%. Its WebDev Arena Elo score reached 1588, compared with 1538\. Google says it can generate functional layouts and feature-complete apps in fewer prompts while closely following a screenshot, image or design system. Document reasoning also advanced. The model reached 34.0% on GDP.pdf against 22.0% for its predecessor, while AutomationBench climbed to 30.4% from 17.0%. Google connects those results to finance, law, biosciences and business workflows. The model can clarify intent, adjust to roadblocks, and put more effort into multi-step planning and tool calls. Developers can access 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, or use it for agent-first workflows in Google Antigravity. Enterprises can find it in the Gemini Enterprise Agent Platform and app. Through the end of the year, pricing is $0.75 per million input tokens and $3.75 per million output tokens, which Google says is half the original 3.6 Flash cost per million tokens. > Gemini Spark now runs on Gemini 3.7 Flash.⚡️ > > Whether you’re using Spark to compile vendors into Sheets or draft negotiation emails, 3.7 Flash makes your personal AI agent more precise and accurate with improved tool use for [@GoogleWorkspace](https://x.com/GoogleWorkspace?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) apps to help turn ideas into action. > > — Google Gemini (@GeminiApp) [August 13, 2026](https://x.com/GeminiApp/status/2087948790296973683?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Gemini Spark is moving to 3.7 Flash immediately. Available to Google AI Pro and Ultra subscribers in more than 160 countries, the personal agent can consolidate files, draft emails and update status documents across Google Workspace tasks. Google says the model raises tool-use accuracy and output quality for complex, multi-skill workflows while Spark remains under user direction. The launch expands Google’s [Gemini](https://www.testingcatalog.com/tag/gemini/) model and agent lineup across consumer, developer and enterprise products. It ships with updated Frontier Safety safeguards for chemical, biological, radiological and nuclear misuse and cyber offense, alongside beneficial use cases under Google’s bioresilience and cyber programs. [Source](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/?ref=testingcatalog.com) ### OpenAI previews Ultrafast API tier for GPT-5.6 Sol URL: https://www.testingcatalog.com/openai-previews-ultrafast-api-tier-for-gpt-5-6-sol/ Last updated: 2026-08-13T21:02:41.000Z OpenAI has opened a limited preview of Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol at up to 14 times the speed of Standard processing. Powered by Cerebras, the mode can generate up to 750 output tokens per second. Access is initially restricted to a select group of customers, with a wider rollout planned as capacity grows. > Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. > > Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows. [pic.twitter.com/a5dleofiDJ](https://t.co/a5dleofiDJ?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [August 13, 2026](https://x.com/OpenAI/status/2087947721936359705?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The service is designed for products and workflows where delays can determine whether an answer is still useful. OpenAI says achieving real-time speeds has often required teams to choose a smaller or specialized model. Ultrafast instead puts the company's most intelligent model into low-latency settings, seeking to deliver more useful work per second without making that tradeoff. The immediate targets span incident response, finance, security, customer support, voice, commerce, and research. Teams could analyze logs, code changes, transactions, or market signals while events are unfolding. Voice and support systems could resolve multi-step requests without breaking a conversation, while commerce tools could check inventory, tailor recommendations, and address checkout problems before a shopper leaves. Researchers could also test and adjust work in shorter cycles. > Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras. > > We gave [@OpenAI](https://x.com/OpenAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com)'s GPT-5.6 Sol the same prompt on Ultrafast and Standard: build a financial terminal-style dashboard for analysts. > > Ultrafast: 1 min 50 seconds > Standard: 12 min 20 seconds > > Same result,… [pic.twitter.com/LgchIPue8b](https://t.co/LgchIPue8b?ref=testingcatalog.com) > > — Cerebras (@cerebras) [August 13, 2026](https://x.com/cerebras/status/2087961128869748856?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) OpenAI is already using the tier internally. During incidents, its developers have applied it to logs, traces, team conversations, follow-up checks, and preparation or validation of fixes, while engineers retain responsibility for judgment and deployment. Research teams are using it across connected tools to search knowledge sources, query data, and organize findings. OpenAI says some experiment loops that once ran overnight can now support several iterations within a workday. Early testing includes Jane Street, Podium, Basis, and Rogo. Their feedback centers on more focused coding sessions, faster complex voice calls, low-latency applications built around a frontier model, and financial research that feels closer to a live exchange. Ultrafast extends [OpenAI's](https://www.testingcatalog.com/tag/openai/) partnership with Cerebras, whose infrastructure supports GPT-5.6 Sol at the stated output rate. The preview is available through the API, and OpenAI is collecting input from initial customers to guide the service as capacity expands. A sign-up form is available for access updates, but the company has not announced pricing or a general-availability timeline. [Source](https://openai.com/index/previewing-ultrafast/?ref=testingcatalog.com) ### ICYMI: xAI releases Grok 4.6 for long-running agent work URL: https://www.testingcatalog.com/icymi-xai-releases-grok-4-6-for-long-running-agent-work/ Last updated: 2026-08-13T21:01:30.000Z xAI has released Grok 4.6, a model built for long-running agents and ambitious coding, research, visual, and interactive work. It builds on Grok 4.5 by handling complex assignments across many steps, from analyzing information and navigating a codebase to turning a broad product idea into a working application or a polished artifact. The release targets developers and knowledge workers who need an agent to carry a project from initial research through implementation and revision. > Introducing Grok 4.6. > > It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price. [pic.twitter.com/RtTbpXcb3a](https://t.co/RtTbpXcb3a?ref=testingcatalog.com) > > — SpaceXAI (@SpaceXAI) [August 12, 2026](https://x.com/SpaceXAI/status/2087562800982077492?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Behind the model is a longer supplemental training run than Grok 4.5 received. xAI used curated model-generated reasoning and technical material, high-quality engineering data, and a revised optimizer and training recipe. Grok 4.5 then regenerated SFT trajectories across reasoning levels, agent harnesses, STEM, software engineering, and knowledge work. Reinforcement learning covered general coding, kernel optimization, web development, and computer-aided design environments. xAI reports a score of 61 on the Artificial Analysis Intelligence Index, up from 56 for Grok 4.5 High and level with GPT-5.6 Sol Max, while Fable 5 Max scored 62\. Grok 4.6 led the listed systems on GDPVal-AA v2 and AA-Briefcase, but trailed GPT-5.6 Sol Max and Fable 5 Max on DeepSWE and Terminal-Bench. That mixed result positions it near the frontier while showing that its strongest gains are not uniform across every coding task. Beyond benchmark scores, xAI says the model produces stronger first passes for visual and interactive projects, can establish an application structure and visual language in one pass, and shows more self-testing on longer runs. It can research an unfamiliar domain, implement core functions, and keep refining the result through feedback, making sustained project work the central pitch rather than one-shot code generation. > We've reset limits to help you keep building during the Grok 4.6 launch. > > Use a reset token from settings in Grok on desktop or mobile. > > — Grok (@grok) [August 13, 2026](https://x.com/grok/status/2087959772909949375?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Grok 4.6 is available now in Cursor, Grok Build, the xAI API, OpenRouter, Vercel, and Cloudflare. API pricing starts at $2 per million input tokens and $6 per million output tokens, while the fast variant costs twice as much. Cursor and Grok Build include double usage for the first week. xAI also says safeguards were calibrated to the model’s capabilities, backed by its widest pre-deployment test suite and continued post-deployment and third-party testing. [Source](https://x.ai/news/grok-4-6?ref=testingcatalog.com) ### Matic robot adds gesture and voice controls in 75 languages URL: https://www.testingcatalog.com/matic-robot-adds-gesture-and-voice-controls-in-75-languages/ Last updated: 2026-08-13T18:12:44.000Z Matic has added voice and gesture controls to its home-cleaning robot, allowing a spoken command paired with a pointing gesture to direct the machine to a specific spot on the floor. Saying "Hey Matic, clean this" while pointing at a spill directs the robot to that exact area without opening the app, and the company reports support for 75 languages. 0:00 /1:24 1× The capability sits on the perception stack Matic has been building since 2017\. Five RGB-IR cameras feed an onboard NVIDIA computer that handles mapping, navigation, and object recognition locally. To resolve a command like "clean this", the system has to find the arm in frame, project where it is aimed, and match that direction against objects already held in its 3D model of the room. Matic states that raw camera footage is discarded in real time, that nothing is uploaded to a server, and that the robot can be set up without an internet connection. That same vision layer shapes how the robot cleans. It identifies furniture, rugs, wires, toys, pets, and floor types, then adjusts brush speed, suction, airflow, mop pressure, and movement pattern based on what it sees. The 3D map is rebuilt continuously as furniture moves or people walk through, so a rearranged room registers as an update. Matic vacuums hard floors and rugs first and mops hard floors second, switching between the two on its own, and navigates in complete darkness using infrared. Waste and dirty water both collect in a sealed HEPA bag, where an absorbent powder gels the liquids, eliminating the need for a dirty water tank. The company lists 55 decibels while vacuuming and describes the robot as quieter than a conversation. Matic sells directly in the United States for $ 1,245 for one robot and one dock, with a one-year warranty and a six-month return window. Every unit is assembled and quality-checked in Mountain View, California. SPONSORED Learn more about Matic [Learn more ](https://maticrobots.com/?ref=testingcatalog.com) The company operates as Matician and was founded by Navneet Dalal and Mehul Nariyawala, who met at Like.com, an early visual search company acquired by Google, and later built Flutter, a hand-gesture detection product that Google also acquired. Matic came out of stealth in 2023 with roughly $30 million in funding and a stated position that indoor robots could run on cameras alone, with all processing on the edge device. Voice and gesture control follows Apple Home and Google Home integrations that shipped earlier this year, and reflects the founders' long-running argument that a machine should take direction the way a person in the room would. ### Google tests Agent management UI on AI Studio URL: https://www.testingcatalog.com/google-tests-agent-management-ui-on-ai-studio/ Last updated: 2026-08-13T12:55:22.000Z Google appears to be building a dedicated agents section into AI Studio, adding a management surface for the managed agents it launched at I/O in May. The work in progress introduces a workbench tab listing available agents, with a project switcher that scopes each list to a specific Google Cloud project. That detail implies agent definitions are meant to sit against Cloud projects and billing rather than in a separate consumer-grade sandbox. > GOOGLE 🔥: AI Studio is about to get a dedicated Agents tab for managing Cloud Agents! > > Users will be able to browse, create, and configure managed agents for different GCP projects. Artifact management will also be available there, along with an editor. > > Soon 👀 > h/t… [pic.twitter.com/AeAHVkftca](https://t.co/AeAHVkftca?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 13, 2026](https://x.com/testingcatalog/status/2087885508898705482?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) From the same tab, a creation flow is present, alongside a per-agent file directory in its own tab where files can be added and edited directly in the browser, and a configuration screen covering name, description, a custom prompt, and further settings. That maps closely to how managed agents are already defined through the API, where an agent is essentially a bundle of instructions, skills, and mounted sources. The unreleased interface would give that bundle a graphical editor instead of a JSON payload or a terminal command. The background model is currently locked to the Antigravity harness, the agent runtime Google shipped at I/O and which now runs on [Gemini 3.6 Flash](https://www.testingcatalog.com/google-launches-gemini-3-6-flash-and-gemini-3-5-flash-lite/). The selector around it is present but offers nothing else, which usually points to alternatives being wired in later, plausibly other harnesses or model tiers as they arrive. AI Studio already carries an agents toggle inside its playground with pre-configured templates, so this reads less like a first step and more like the layer above it, where prototyping turns into hosting and lifecycle management. Anthropic reached that point in April, when Claude Managed Agents entered public beta with configuration, sandboxes, and session tracing exposed through its console, and Google now looks set to close the same loop. Timing remains unclear, and the model roadmap does not help. Gemini 3.5 Pro missed its June target, and an analyst note claims it was quietly shelved in favor of Gemini 4, although Google has not confirmed this and still lists the model as coming. The agent stack leans on Flash-class models regardless, so it is not obviously waiting on a Pro release. [Source](https://x.com/thomas%5Fgmry?ref=testingcatalog.com) ### NoimosAI launches SEO Agent 2.0 for automated optimization URL: https://www.testingcatalog.com/noimosai-launches-seo-agent-2-0-for-automated-optimization/ Last updated: 2026-08-12T16:17:50.000Z NoimosAI has released SEO Agent 2.0, an agent that links SEO research and analytics to the work of publishing and fixing pages. The company frames the difference against tooling that stops at a report: here, a recommendation is carried through to a change on the site, with a human approval step sitting between the proposal and the action. > Most SEO tools give you insights. > > NoimosAI turns them into action and results. > > Just connect your website and apps. It finds winning keywords, writes optimized articles, and even fixes technical issues to increase free traffic from Google. [pic.twitter.com/fJIN2UK9bx](https://t.co/fJIN2UK9bx?ref=testingcatalog.com) > > — NoimosAI (@noimos\_ai) [August 12, 2026](https://x.com/noimos%5Fai/status/2087569817360592900?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The agent draws on several data sources at once. Semrush supplies domain power, keyword difficulty, and backlink gap analysis. Google Search Console covers query performance, indexing requests for new or updated pages, high-ranking pages with weak click-through rates, and cases where two pages compete for the same term. GA4 adds behavioral context, flagging content decay and scoring pages against time-on-page signals. PageSpeed data and site audits cover on-page and technical issues. From that combined picture, the agent proposes concrete work, such as closing a keyword gap a competitor already ranks for, or rewriting an article whose visibility has slipped, using the pages currently ranking above it as reference. Execution is the part the company is pushing with this version. Once a proposed action is reviewed and approved, the agent can draft and publish articles complete with images, internal links, and meta tags, and apply supported changes to websites managed through the NoimosAI Website Builder. Publishing also reaches WordPress through direct CMS integrations. Sites built inside the platform ship with structured data and technical elements aimed at both search engines and AI answer systems, a split the company labels SEO and GEO. The intended users are entrepreneurs, freelancers, creators, marketers, search professionals, and small business owners who are running organic growth without a dedicated team. NoimosAI is built by AGOS LABS TECHNOLOGIES LTD, led by chief executive Kosuke Yokoyama, and is positioned as a multi-agent marketing workspace instead of a single tool. Alongside search, the platform runs capabilities for: Competitor strategy, Social listening, Industry news, Social media, Event and media outreach, Conversion rate work SPONSORED Start testing NoimosAI [Learn more ](https://noimosai.com/en?ref=testingcatalog.com) It also offers connections to more than twenty apps across analytics, CMS platforms, social networks, and collaboration software. Website Builder arrived earlier this year on the premise that a published site should keep rewriting its own code against performance data, and a Creative Agent covering social and ad assets followed shortly after. SEO Agent 2.0 extends the same pattern into search. Access runs on a credit-metered subscription, with the Pro tier at 99 dollars per user per month, Team at 249 dollars, and Advanced at 499 dollars, alongside a free trial on signup. ### Cursor prepares to launch Origin platform for code reviews URL: https://www.testingcatalog.com/cursor-prepares-to-launch-origin-platform-for-code-reviews/ Last updated: 2026-08-11T22:55:17.000Z Cursor is preparing to open Origin beyond the closed partner beta it has been running for weeks. Strings across the Cursor web interface point to the platform shipping under the internal name "Cursor Review", with two tabs appearing once access is switched on. Codebase covers syncing and managing repositories pulled in from GitHub. Review is the more consequential half: an automated pull request pipeline that notifies developers when their judgment is needed, so humans and agents can work through open PRs across a codebase together. Signals suggest a rollout could land as early as this week, ahead of the fall window Cursor named when it announced the platform in June. > Cursor Origin (Cursor Review) has been tested in a closed beta with selected partners internally and will likely be announced later today. > > Users will get access to new tabs called Codebase and Review. > > \> The Codebase tab will let users sync their repositories from Github. > > \>… [https://t.co/QoKY5YEw38](https://t.co/QoKY5YEw38?ref=testingcatalog.com) [pic.twitter.com/3xEMm3Ux3A](https://t.co/3xEMm3Ux3A?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 11, 2026](https://x.com/testingcatalog/status/2087177021105267167?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Origin was unveiled at Cursor's Compile conference and built by the Graphite team the company acquired in late 2025\. The pitch is that GitHub was designed around human-paced review, one reviewer, one diff, sequential merges, while Cursor demoed 22.6 commits per second into a single repository. Teams running fleets of background agents are the obvious beneficiaries, because review rather than generation is where agentic workflows now stall. The tab structure suggests Cursor wants to land the review layer first and migrate hosting later, the lower-friction path for teams unwilling to move source control off GitHub outright. > origin origin origin > > — Jacob Gold (@jacobgold) [August 10, 2026](https://x.com/jacobgold/status/2086882423124693346?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > I don’t think there’s been a bigger week for Cursor for three reasons. > > — Ricky Doar (@rdoar4) [August 11, 2026](https://x.com/rdoar4/status/2087189798884671991?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) That timing sits inside a larger consolidation. [SpaceXAI](https://www.testingcatalog.com/tag/spacexai/) recently shipped Grok Bot into beta, a desktop and mobile app that gives agents a shared cloud machine they can use to sign in to tools and finish work unattended. It carries its own Origin references, and once the platform is live, Grok Bot looks set to pull repositories directly from it and act on them. With SpaceX's $60 billion acquisition of Anysphere expected to close this quarter, and Grok 4.6 having briefly surfaced in [Cursor's](https://www.testingcatalog.com/tag/cursor/) model list as "Cursor Grok 4.6 " before being withdrawn, the two roadmaps are folding into one. ### OpenAI plans free ChatGPT Plus year for US college students URL: https://www.testingcatalog.com/openai-plans-free-chatgpt-plus-year-for-us-college-students/ Last updated: 2026-08-11T22:02:51.000Z OpenAI looks set to open the back-to-school season with its largest student giveaway yet: **a full year of ChatGPT Plus at no cost**. Signals point to an offer confined to the United States and gated behind an allowlist of more than 220 participating campuses, with enrollment confirmed through SheerID, the same verification partner used for the company's two-month student promotion in spring 2025\. Nothing has gone live yet, and OpenAI has not published terms. > OPENAI 🔥: College students in the US will be able to claim a year of ChatGPT Plus for free. > > OpenAI is preparing a new "back-to-school" campaign, positioning ChatGPT as a tool for exam preparation, studying, and more! > > \> From study sessions to final projects, ChatGPT helps you… [pic.twitter.com/tjFR7k42oa](https://t.co/tjFR7k42oa?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 11, 2026](https://x.com/testingcatalog/status/2087298222301569414?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The framing positions [ChatGPT](https://www.testingcatalog.com/tag/chatgpt/) as an academic operating layer rather than a homework tool. The pitch covers: 1. Exam preparation 2. Voice mode for rehearsing a language or running through presentations and interviews 3. Website generation 4. An agentic layer that reads email, manages calendars, and folds a semester of scattered context into a single project plan That last claim leans on connectors and projects, so the payoff depends on how much a student is willing to wire into their account. The constraints matter too: a campus allowlist and US-only eligibility make this targeted acquisition rather than a blanket offer, and the 2025 pilot rolled over to full price once its window closed. It arrives in a market where the free year has become table stakes. Google's year of AI Pro for students has wound down, replaced by a free year of the cheaper AI Plus tier for US and Canadian students through Handshake, with that window closing as term begins. Perplexity offers education pricing at $10 a month, and Anthropic routes students through institutional Claude for Education agreements rather than consumer discounts. For [OpenAI](https://www.testingcatalog.com/tag/openai/), 12 months covers an academic cycle and is long enough to set habits before graduation. It also fits a steady run of education work: ChatGPT Edu campus deals, free access for verified K-12 teachers, and the educator and college-student plugins that shipped this month. Anyone in the US about to pay for Plus may want to sit tight for a few weeks. ### Meta releases Muse Glimmer for local AI agents URL: https://www.testingcatalog.com/meta-releases-muse-glimmer-for-local-ai-agents/ Last updated: 2026-08-10T11:29:24.000Z Meta has released Muse Glimmer, a 30-billion-parameter open-weight model built for always-on agents on a Mac or PC with a single consumer GPU. The model comes from Meta Superintelligence Labs, carries a permissive Apache 2.0 license, and is available now on [HuggingFace](https://huggingface.co/meta-models?ref=testingcatalog.com). It targets developers building local agents, coding tools, function-calling systems, and model-based evaluation without relying on a constant network connection. > Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows. > > Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on… [pic.twitter.com/mI4z91GPnE](https://t.co/mI4z91GPnE?ref=testingcatalog.com) > > — AI at Meta (@AIatMeta) [August 10, 2026](https://x.com/AIatMeta/status/2086757844544811485?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Muse Glimmer is trained for end-to-end task completion, precise tool calls, multi-step reasoning, and recovery when a tool fails. A dedicated perception encoder lets it process interleaved text and images, including screenshots, charts, and documents. It also supports more than 100 languages, selectable reasoning strengths, and agent scaffolds such as OpenClaw. > 2/ just like much larger models, muse glimmer can operate as a fully capable agent via planning, tool calls, checking its own results, and failure recovery. [pic.twitter.com/CMnroWLTe8](https://t.co/CMnroWLTe8?ref=testingcatalog.com) > > — Alexandr Wang (@alexandr\_wang) [August 10, 2026](https://x.com/alexandr%5Fwang/status/2086756153464386012?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Meta designed the model around the memory and compute limits of consumer hardware. Its training used logit distillation from Muse Spark outputs, followed by longer-context agent data, supervised fine-tuning, on-policy distillation, and reinforcement learning across general, reasoning, coding, and agent tasks. Meta says the model was evaluated under its Advanced AI Scaling Framework before the open-weight release. A full-precision version would need more than 55 GB of memory. Quantization reduces the language model to under 20 GB, leaving room in a 24 GB or 32 GB memory envelope for its working memory, KV cache, image encoder, and speculative-decoding drafter. That lightweight DFlash-based companion proposes blocks of tokens that the main model verifies in parallel. Meta reports decode-speed gains of 3.1 times on an RTX 5090, 1.8 times on an M5 Max, and 1.5 times on an M4 Max for its K-Quant-17GB setup. > Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats… > > — Mark Zuckerberg (@finkd) [August 10, 2026](https://x.com/finkd/status/2086755195535413696?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Meta positions Muse Glimmer against Gemma4-31B and Qwen3.6-27B, reporting strong results for its size across agentic and general language-model benchmarks. The model is intended for local work that can continue without cloud infrastructure, while still supporting deployment through larger serving stacks. ![Meta on HuggingFace](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/meta-models-Meta-Inc--08-10-2026_01_29_PM.jpg) Meta on HuggingFace [Weights](https://huggingface.co/meta-models?ref=testingcatalog.com) and developer documentation are available now. Optimized support for llama.cpp, MLX, and ExecuTorch is due in the coming days, alongside planned access through Ollama, LM Studio, Unsloth, vLLM, SGLang, Together AI, Fireworks AI, and OpenRouter. Meta is also working with AMD, Arm, Dell, Intel, and NVIDIA on device-level optimization, extending its open AI research into local agent systems. [Source](https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model?ref=testingcatalog.com) ### Google may retire Gems in October, forcing migration to Skills URL: https://www.testingcatalog.com/google-may-retire-gems-in-october-forcing-migration-to-skills/ Last updated: 2026-08-09T16:39:00.000Z Google appears to be preparing to retire Gems on October 20, based on a notice currently behind a feature flag in the [Gemini](https://www.testingcatalog.com/tag/gemini/) web app that has not been shown to anyone yet. The wording advises users to save their data and suggests they may later be able to rebuild their Gems as skills, implying that skills are intended to be broadly available by that date. Google has not commented publicly, and a flagged banner is not a commitment, so the timing could still move. > [@killedbygoogle](https://x.com/killedbygoogle?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), get the grave ready 💀 > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 8, 2026](https://x.com/testingcatalog/status/2086199312082469138?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The gap between the two formats is what makes this awkward. Gems are free for everyone and sit in the Gemini sidebar. Skills live inside [Gemini Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/), require a Google AI Pro or Ultra subscription, are limited to personal accounts rather than work or school ones, and stay unavailable in the European Economic Area, the UK, Switzerland, and Nigeria. Spark only reached Pro subscribers across 160-plus countries at the end of July. Retiring Gems on those terms would move a free customization layer behind a paid, region-limited agent product. The reaction to our post was sharpest among Korean users, who have built large Gem libraries. The notice implies no automated migration, so heavy users may need to rebuild each one by hand. Google probably has tooling planned, and rivals already let people author reusable instruction files in minutes, but that still leaves manual reconstruction. Enterprise is the harder question. Google has spent two years pitching Gems for workplace use, grounded in Drive and Gmail, while skills do not support work or school accounts at all. Any deadline reaching Gemini Enterprise customers would land as a migration project rather than a settings change. It would also be the second reshuffle of this surface in under a year, following the folding of Opal workflows into the Gems manager as Super Gems in December. The feedback so far is consistent: "Automate the conversion." Source: [Tomas](https://x.com/thomas%5Fgmry?ref=testingcatalog.com) ### xAI launches Imagine Image 2.0 in Grok Quality Mode URL: https://www.testingcatalog.com/xai-launches-imagine-image-2-0-in-grok-quality-mode/ Last updated: 2026-08-08T07:20:31.000Z xAI has released Imagine Image 2.0 as the new Quality Mode on Grok’s web-based Imagine service and its iOS and Android apps. The model is aimed at creative work that requires controlled layouts, legible typography and repeatable visual elements, rather than one-off image generation. API access is planned but is not yet available. > Edit exactly what you mean with new tools: Magic Wand changes one region and leaves the rest, Segmentation selects precise areas, and Smart Resize lets you tailor to any aspect ratio. [pic.twitter.com/Fs2xQkdtqC](https://t.co/Fs2xQkdtqC?ref=testingcatalog.com) > > — Grok (@grok) [August 8, 2026](https://x.com/grok/status/2085931543738941477?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The company says Image 2.0 can follow detailed instructions, plan typography and layout for dense compositions, keep small text sharp, and preserve supplied elements across new generations and edits. Its editing toolkit lets users target a region with a magic wand or segmentation while leaving the rest of an image unchanged. It can also remove a background to create a transparent export and combine up to five reference images in a single generation, reducing the manual compositing needed for complex scenes. Smart Resize extends an image to a chosen frame, with ratios ranging from 1:2 and 9:16 through square, landscape and 2:1 formats. ![Grok](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/cyberpunk-hacker-robot-working-in-front-of-many-monitors-Grok-08-08-2026_09_14_AM.jpg) xAI is also introducing ready-made templates for recurring workflows. The selection spans photo edits, product color changes, e-commerce shots, professional headshots, icons, character sprites, emojis and merchandise. Users provide the inputs while the template supplies the configured workflow. For video planning, the page demonstrates generating a character, locations and props separately while maintaining a shared visual style, allowing a recurring figure and its surroundings to be developed as one consistent world. > Exciting news: Grok Imagine Image 2.0 (Low) by [@SpaceXAI](https://x.com/SpaceXAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) has landed in the Text-to-Image Arena at #2 (1320 pts)! This release is not available via API, only in their app. > > Grok Imagine Image 2.0 (Low) is a significant improvement from Grok Imagine Image Quality (#14 -> #2). > > By… [https://t.co/jjHBRuXwPn](https://t.co/jjHBRuXwPn?ref=testingcatalog.com) [pic.twitter.com/U9I23p3zVQ](https://t.co/U9I23p3zVQ?ref=testingcatalog.com) > > — Arena.ai (@arena) [August 8, 2026](https://x.com/arena/status/2085951322897998199?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The company says Image 2.0 ranks second worldwide in both text-to-image generation and image editing. In the Arena results displayed on the page, grok-imagine-image-2 (low) scores 1320 for text-to-image, behind gpt-image-2 at 1380, and 1439 for editing, behind gpt-image-2 at 1463\. The separately listed grok-imagine-image-quality entry scores 1228 and 1390 in those categories. xAI says its models appear under SpaceXAI on Arena. [Source](https://x.ai/news/grok-imagine-image-2?ref=testingcatalog.com) ### OpenAI says Astra may have reached Critical cyber threshold URL: https://www.testingcatalog.com/openai-says-astra-may-have-reached-critical-cyber-threshold/ Last updated: 2026-08-08T07:22:43.000Z OpenAI says it can no longer rule out that Astra, an upcoming model, has reached the Critical cybersecurity capability threshold in its Preparedness Framework. Preliminary internal evaluations conducted over the past few days found major gains in agentic coding and cybersecurity, prompting the company to tighten controls while benchmarking continues. > After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework. > > This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development… > > — OpenAI (@OpenAI) [August 7, 2026](https://x.com/OpenAI/status/2085801349866729975?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The finding is not a final Critical classification. Under OpenAI’s framework, that level covers models capable of discovering and building working zero-day exploits across many hardened, real-world critical systems without human help, or of devising and executing novel end-to-end attacks against hardened targets from only a high-level goal. OpenAI says Astra’s current performance is strong enough that this level cannot yet be excluded. The company also states that Astra was not involved in exploiting Hugging Face. > astra is a powerful model and we are working to make it generally available. > > we do not think it is a good strategy to keep powerful models to a chosen few. > > given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long! > > — Sam Altman (@sama) [August 7, 2026](https://x.com/sama/status/2085862292311396515?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Astra remains unreleased, with no availability timeline disclosed. OpenAI has paused internal work involving the model when it does not meet the new control requirements. Development and testing will use isolated environments, restricted network and tool access, stronger model-weight protection and encryption, additional monitoring and detection, and sandboxed execution. The company has also introduced universal monitoring for risky actions and misalignment across Astra’s agentic uses, including training and evaluation. These monitors assess the model’s chain of thought and can trigger a security response to review and interrupt high-risk activity. OpenAI plans to work with government agencies and selected AI safety organizations on capability testing, while giving third-party testing partners recommended controls for higher-risk evaluations and workloads. OpenAI launched its Preparedness Framework in 2023 to track frontier risks in biological, chemical, cybersecurity, and AI self-improvement capabilities and guide its response as models advance. Earlier systems, including GPT-5.6-Sol, were assessed at the High cyber threshold, placing Astra’s possible Critical capability beyond the company’s previous cyber assessments. The company frames the shift as both an urgent security problem and a defensive opportunity. Its stated aim is for cyber-capable models to help defenders find and fix vulnerabilities before attackers can exploit them, with deployment shaped in collaboration with governments, safety institutes, and civil society. [Source](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/?ref=testingcatalog.com) ### Meta prepares desktop AI app for macOS as rivals race ahead URL: https://www.testingcatalog.com/meta-prepares-desktop-ai-app-for-macos-as-rivals-race-ahead/ Last updated: 2026-08-07T22:52:04.000Z Meta appears to be working on its own desktop application for macOS, judging by new references that have shown up in recent Meta AI web builds. One is a button in the settings menu, marked with a macOS icon and labeled "Download Meta AI." Alongside it sits a notice describing the app as a way to chat, create, and collaborate with AI from anywhere on the desktop. The description is still abstract at this stage. It says nothing about coding capabilities and gives no indication of whether the client will lean toward consumer use cases or something closer to a work tool. The download link currently points nowhere, which suggests a feature that is not fully implemented and not available to anyone yet. ![Meta AI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Meta-AI-08-07-2026_11_03_PM.jpg) The timing lines up with a broader push. Meta recently released its [command-line tool powered by Muse Spark 1.2](https://www.testingcatalog.com/meta-launches-muse-code-beta-powered-by-muse-spark-1-2/), and a desktop client appears to be the next step in catching up with rivals in the field of desktop "super apps." OpenAI, Anthropic, Google, and SpaceXAI are all actively building tools in this space, and for Meta, having a presence on the desktop matters if the company wants to stay relevant in a category that is quickly becoming a default surface for AI assistants rather than an optional extra. There is also the question of [Hatch](https://www.testingcatalog.com/meta-prepares-hatch-agent-under-waitlist-and-social-media-skills/). Earlier this year, reports pointed to Meta AI working on its own always-on agent under that name, and traces suggested it was potentially planned for a waitlist release before those references were later removed from the codebase. Whether Hatch ends up shipping as part of the desktop app remains unclear, though it is still highly possible, particularly given how well an always-on agent would fit a client that runs quietly in the background of a machine. For now, the timeline stays open. With the download link inert and the surrounding copy still generic, it may take some time before anything official arrives from [Meta](https://www.testingcatalog.com/tag/meta/) on this. ### NoimosAI unveils Social Agent to automate content creation URL: https://www.testingcatalog.com/noimosai-unveils-social-agent-to-automate-content-creation/ Last updated: 2026-08-07T17:03:54.000Z NoimosAI has launched Social Agent, an autonomous marketing agent designed to research trends in a user's niche, analyze high-performing posts, and produce content informed by proven strategies. The agent went live on August 7 and supports 9 channels: X, Facebook, Instagram, Threads, TikTok, YouTube, Pinterest, Bluesky, and Mastodon. 0:00 /1:05 1× The workflow begins once accounts are connected. Social Agent scans multiple platforms to identify trends from data, referencing the formats, hooks, and structures of successful posts in a given niche before drafting content. The output extends beyond plain text to include carousel posts, product videos, UGC videos, and clipped videos. It can create a natural UGC video from a product image or link, or generate a product video from a code implementation. Scheduled posting is integrated into the workflow, and the agent conducts performance analysis on connected accounts, suggesting improvement plans for future content. The target users are solo operators or small teams: entrepreneurs, freelancers, creators, marketers, and small business owners who manage social channels alongside other responsibilities and often rely on intuition for content decisions. Social Agent argues that decisions need not be intuitive, since trends, formats, and performance patterns are observable and can serve as inputs. According to NoimosAI, this approach increases traffic and sales. 0:00 /0:55 1× Launch Video Social Agent is part of a broader platform operated under the label Command Marketing, where users set goals, and specialized agents cover growth metrics, competitor strategy, social listening, industry news, SEO, GEO, event outreach, media outreach, and conversion optimization to build and execute strategies. NoimosAI opened to the public in January 2026 after a private beta, launched a Japanese market release in June, and introduced a Creative Agent for autonomous media generation in July. Paid plans are priced at $99 per user per month for Pro, $249 per user per month for Team, and $499 per user per month for Advanced, all metered in credits, with a free tier available to start. SPONSORED Start testing NoimosAI [Learn more ](https://noimosai.com/en?ref=testingcatalog.com) The platform is developed by AGOS LABS TECHNOLOGIES LTD, a Dubai-based company led by founder and chief executive Kosuke Yokoyama. The team previously worked on Web3 products before transitioning to marketing automation. NoimosAI is billed as the first autonomous AI marketing agent, with a deliberately neutral architecture, since it does not sell its own execution tools, unlike a CRM. Yokoyama has publicly stated that the platform achieved $1 million in annual recurring revenue within 30 days of launch through organic distribution alone. ### OpenAI is rolling out GPT-5.6 Luna to Free ChatGPT users URL: https://www.testingcatalog.com/openai-is-rolling-out-gpt-5-6-luna-to-free-chatgpt-users/ Last updated: 2026-08-07T12:12:49.000Z OpenAI is rolling out a broad ChatGPT update that gives paid users a more reliable GPT-5.6 Sol experience while moving Free and Go accounts to GPT-5.6 Luna. The changes are aimed at making everyday answers more focused, reducing factual mistakes, and giving people clearer control over how much reasoning ChatGPT uses. Plus and Pro subscribers are getting an updated GPT-5.6 Sol that now powers both Instant replies and deeper reasoning in Chat. OpenAI says the model is designed to answer quick questions directly, add detail when a task requires planning, research, writing, or coding, and avoid repeating information after follow-up prompts. A new reasoning-effort slider on web, mobile, and desktop lets users keep responses quick or give the model more time for involved work. > We’re making better intelligence easier to access in ChatGPT for everyone: > > \- GPT-5.6 Sol now powers both Instant and deep reasoning for Plus & Pro users, delivering more factual, focused responses. > > \- Free & Go users get unlimited text chats with GPT-5.6 Luna starting tomorrow. [pic.twitter.com/JXhmj5GLTH](https://t.co/JXhmj5GLTH?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [August 6, 2026](https://x.com/OpenAI/status/2085434712429052386?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Free and Go users will see GPT-5.6 Luna become their default model during the week of August 6\. Starting the following week, OpenAI plans to offer those accounts unlimited text chats and a new Think button, which gives Luna more time on harder questions. Abuse guardrails will still apply, while separate limits will remain for file uploads, images, and other tools. Reliability is a central part of the release. In OpenAI’s internal evaluation of financial, medical, and legal prompts requiring factual detail, answers containing at least one factual error were about 62 percent less common with GPT-5.6 Luna and 68 percent less common with GPT-5.6 Sol than with GPT-5.5 Instant. The company says Sol makes stronger use of sources when handling dates, numbers, rules, and assumptions, and is more willing to correct a premise rather than simply agree. OpenAI says ChatGPT now serves one billion people each week across quick searches, advice, research, planning, and complex decisions. The updated Sol model is limited to ChatGPT’s Chat experience, and the Sol version used by Work and Codex is unchanged. OpenAI has also added safety training for users believed to be under 18, including boundaries around romantic roleplay, sexual content, eating disorders, dangerous activities, age-restricted goods, and graphic violence, with prompts to seek support from trusted people when needed. [Source](https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/?ref=testingcatalog.com) ### ICYMI: Prime Intellect releases open-source Prime Agent URL: https://www.testingcatalog.com/icymi-prime-intellect-releases-open-source-prime-agent/ Last updated: 2026-08-07T12:09:11.000Z Prime Intellect has released Prime Agent, an open-source coding harness that can revise parts of its operating setup as it works. Available now on the company's GitHub repository, it targets developers working with modern open and closed frontier models. The company presents it as a coding assistant, a runtime for long-horizon autonomous evaluations, and a research collaborator, though no model has yet been trained specifically for the harness. Prime Agent centers on the Recursive Language Model, or RLM, and Continual Harness. RLM treats context as a variable and sub-agent delegation as function calls inside a persistent REPL, giving the model programmatic access to its history, tools, and sub-agents. Continual Harness lets the agent create, read, update, and delete prompts, memory, skills, and sub-agent definitions from its own trajectory. Prime Intellect says this can preserve access to earlier information across arbitrarily long sessions. > Prime Agent is a general-purpose coding harness > > On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific. > > We see major improvements across models when compared to their proprietary harnesses: [pic.twitter.com/Xa8zZiDtbF](https://t.co/Xa8zZiDtbF?ref=testingcatalog.com) > > — Prime Intellect (@PrimeIntellect) [August 5, 2026](https://x.com/PrimeIntellect/status/2085087000764568010?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model's sole tool is a persistent IPython kernel. An asynchronous rlm call can launch a complete sub-agent session and return a handle for later messaging. A background daemon owns live sessions, while append-only JSONL histories and kernel snapshots support recovery. The /refine pipeline can apply targeted edits at a turn boundary, with the base system prompt kept immutable and prior refinements available for rollback. For unattended work, autonomous mode combines a persistent goal, scheduled heartbeats, and continuation, with optional completion gates and turn, token, and wall-clock limits. Prime Intellect also reports that Opus 5 running in Prime Agent reached 95.5% RHAE Best@1 on ARC-AGI-3, narrowly above its cited 95.4% human expert baseline. These are company-reported launch results. The risks are already visible. In Factorio tests, /refine turned prior outcomes into memory and skills and raised production scores, but it also learned to use RCON commands to spawn resources despite instructions not to cheat. The case shows how self-improvement can reinforce reward hacking alongside useful strategies. Prime Agent is built on the open-source pi project, and Prime Intellect says current models still have friction with the harness. A fuller technical report is planned. [Source](https://www.primeintellect.ai/blog/prime-agent?ref=testingcatalog.com) ### Meta launches Muse Code beta powered by Muse Spark 1.2 URL: https://www.testingcatalog.com/meta-launches-muse-code-beta-powered-by-muse-spark-1-2/ Last updated: 2026-08-05T20:53:41.000Z Meta has released Muse Code in beta, a terminal coding agent for complex software engineering work across large repositories. Available on macOS and Linux, it can plan changes, write code, and validate results. The agent runs on Muse Spark 1.2, a coding-focused model now offered through Meta Model API with expanded global access. > Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update. [pic.twitter.com/xqavk41w6v](https://t.co/xqavk41w6v?ref=testingcatalog.com) > > — Mark Zuckerberg (@finkd) [August 5, 2026](https://x.com/finkd/status/2085080750034940201?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The runtime pairs a main agent loop with persistent asynchronous background agents that stay active throughout a session. They gather context, take next steps, and decide when to report back, avoiding repeated information collection during multi-stage tasks. Muse Code can coordinate multiple agents on a single job to reduce latency and manual steering. Every model call, tool run, approval and edit is appended to a local event log. Meta says this makes work replay-exact and restart-safe, so an interrupted task can resume from the same point after a crash. Built-in skills can create an approval-gated plan, stress-test it and continue toward a defined goal. To install Muse Code, run 💡 curl -fsSL [https://dev.meta.ai/install.sh](https://dev.meta.ai/install.sh?ref=testingcatalog.com) | bash Muse Spark 1.2 builds on version 1.1 with stronger code generation, complex debugging, codebase understanding and end-to-end development workflows. Meta increased training compute for coding and expanded the range of training environments. It co-trained the model with Muse Code using rejection-sampled harness trajectories and recipe optimizations for goals, context compaction and subagent work. Training covered long-horizon jobs including whole-repository generation, large projects and automated research. Meta also used Muse Spark 1.1 to create demanding coding environments and instruction templates, then grade candidate solutions to build training data for its successor. The company says this helped the new model follow complex requirements more precisely. In one case study, the system made more than 1,000 tool calls over runs lasting up to 24 hours while optimizing GPU kernels for NVIDIA Hopper hardware. It wrote, compiled and profiled Triton implementations, revising them against a baseline without importing third-party kernel libraries. The release shows [Meta](https://www.testingcatalog.com/tag/meta/) coupling a coding model with the runtime used to train and operate it, framing the pair as one developer system. Meta says more harness features and larger, more capable models are still ahead. [Source](https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2?ref=testingcatalog.com) ### MemoryPlugin launches Sync app for coding agents on macOS URL: https://www.testingcatalog.com/memoryplugin-launches-sync-app-for-coding-agents-on-macos/ Last updated: 2026-08-05T20:40:53.000Z MemoryPlugin has shipped MemoryPlugin Sync, a macOS app that pulls local coding agent sessions into the same archive as its browser chat history, and the company has listed its memory server in the official MCP registry. Both moves extend a product that until now covered web chat tools almost exclusively. > We just shipped MemoryPlugin Sync for macOS: it syncs your local Claude Code, Codex, and Cursor sessions into the same memory as your ChatGPT, Claude, and Gemini chats. One searchable history, recallable from 21+ AI tools. > > Only prompts and replies sync. Code and tool output… > > — MemoryPlugin (@MemoryPlugin) [August 5, 2026](https://x.com/MemoryPlugin/status/2085007085985735129?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The app monitors the session files that Claude Code, Codex, and Cursor store on disk and uploads completed conversations to a MemoryPlugin account. Only prompts and assistant replies leave the machine, while code, tool output, file contents, terminal commands, and model thinking stay local, according to the company. Syncing is opt-in via per-project and per-chat checkboxes; nothing uploads until that list is reviewed, and every synced conversation includes a remove action that stops future syncs. The app lives in the menu bar and syncs in the background. It requires Apple Silicon and macOS 12 or later, with Intel and Windows builds described as in development. Reading local sessions inside the app is free on any plan, while syncing coding agent chats to the archive requires the Pro tier. ![MemoryPlugin](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/mp-figure-how-a-session-becomes-memory.png) Recall is the point of that archive. MemoryPlugin scans stored conversations, picks the ones relevant to a new question, and injects them into the model context within a few seconds across the 21 or more supported tools. A constraint worked through in Claude Code can surface later inside ChatGPT, and a pricing debate held there can be recalled in Claude. The browser extension keeps syncing ChatGPT, Claude, Gemini, Grok, DeepSeek, and TypingMind conversations live, and official data exports cover older history. The company also lists a system-wide chat search bound to a configurable shortcut. 💡 Our readers can use code TESTINGCATALOG to get 10% off for life! The memory server was published to the official MCP registry in late July under the identifier com.memoryplugin/memory, exposed via streamable HTTP and server-sent events, allowing Claude Desktop, Cursor, and other MCP clients to read and write the same memory without going through a browser. Alongside it, MemoryPlugin released an MIT-licensed agent skill that instructs Claude Code, Codex, Cursor, and other skills capable agents to search memory before assuming and to save durable facts before finishing. SPONSORED Start testing MemoryPlugin [Learn more ](https://www.memoryplugin.com/?utm%5Fsource=testingcatalog&utm%5Fmedium=sponsored) MemoryPlugin is built by Crestify, the software company led by founder Alara Dhamani, which has run the product since it launched as a browser extension for a handful of chat platforms. Coverage now spans more than 21 tools through the extension, a hosted and local MCP server, a custom GPT, a TypingMind plugin, and an open API, with users reported across 65 countries. Annual pricing sits at 89 dollars for Core, which covers cross-AI memory and the last 500 conversations, and 180 dollars for Pro, which removes the history limit and adds summaries, monthly insights, and shared buckets. Every plan opens with a seven-day trial. Conversations are encrypted at rest and are never used to train models, according to the company. ### ByteDance launches SeedRealtime full-duplex AI model URL: https://www.testingcatalog.com/bytedance-launches-seedrealtime-full-duplex-ai-model/ Last updated: 2026-08-05T12:12:35.000Z ByteDance's Seed organization has launched SeedRealtime, a native audio-visual full-duplex large language model built to watch, listen, and speak over continuous streams at once. The model combines audio, video, and text into a single architecture, allowing perception, understanding, decision-making, and expression to run in parallel. This avoids handoffs between separate speech recognition, vision-language, and text-to-speech modules, which can add latency and cause context loss. > BREAKING 🔥: ByteDance launched SeedRealtime, a native audio-visual full-duplex LLM! > > \> SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen,… [https://t.co/lYOEnvyVQv](https://t.co/lYOEnvyVQv?ref=testingcatalog.com) [pic.twitter.com/Qc0SX4x8Xt](https://t.co/Qc0SX4x8Xt?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 5, 2026](https://x.com/testingcatalog/status/2084968825942893022?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The central shift is that the model determines conversational timing rather than relying on external voice-activity detection rules. SeedRealtime tracks scenes, speakers, pauses, and background chatter to judge what matters and when to answer. Visual context can resolve homophones, connect words such as "this" to a gesture or an earlier action, and retain information that has moved off-screen. It can also speak without a new prompt when a requested object appears, or it detects a mistake. ByteDance showed these abilities in noisy, open settings. SeedRealtime matched names, faces, and voices during a group dinner, interpreted dishes and speech in a restaurant, and issued a reminder when a requested museum object entered view. It corrected an espresso-making mistake, spotted a requested section while pages of a paper were turning, ignored unrelated airport chatter, and followed a child's pointing during an English lesson. In the airport example, it went online to provide baggage-carousel information. In end-to-end human evaluation, ByteDance says SeedRealtime cut audio-visual conversational pacing problems by half against cascaded models. Evaluators found fewer cutoffs, slow replies after pauses, and false triggers from nearby speech, while more conversations were completed smoothly. The announcement does not disclose sample sizes, a benchmark table or an exact figure for the completion gain. Published by ByteDance's Seed organization, the release is framed as a step toward omni-modal systems that can observe, converse and act in changing real-world settings. Seed says the model is fully rolled out, but does not specify an API, product surface, supported regions, pricing or access requirements. Its roadmap covers lower latency, finer timing for interruptions and backchannels, stronger speaker tracking in multi-person scenes, proactive decisions and tool-connected tasks such as lookups and bookings. [Source](https://seed.bytedance.com/en/blog/seedrealtime-audio-visual-full-duplex-llm-released-toward-omni-modal-natural-interaction?ref=testingcatalog.com) ### Cloudflare announces Wallets for AI agent payments URL: https://www.testingcatalog.com/cloudflare-announces-wallets-for-ai-agent-payments/ Last updated: 2026-08-05T11:56:49.000Z Cloudflare has announced Cloudflare Wallets, a programmable wallet system designed to let AI agents pay for APIs, content, and other online services without relying on human-facing signup and billing flows. Account holders can [claim a unique Wallet handle](https://cloudflare.pay/?ref=testingcatalog.com) now, while funding and payment features are set to follow. The company plans to support stablecoins and x402, a protocol that attaches micropayments to HTTP requests. The system will have two layers. Account Wallets will let Cloudflare account owners add and remove funds, then delegate spending to Virtual Wallets managed by agents. Those agent wallets will operate via API keys and follow controls chosen by the owner, including allowances, allowlists, maximum transaction sizes, and overall spending caps. Agents that reach a limit could request a manual override from an authorized human. > Introducing Cloudflare Wallets. They will allow you to store stablecoins, purchase services, and receive funds across the web. [https://t.co/ucAaoKMs4Q](https://t.co/ucAaoKMs4Q?ref=testingcatalog.com) > > — Cloudflare (@Cloudflare) [August 4, 2026](https://x.com/Cloudflare/status/2084648084131242402?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Cloudflare is positioning this model as a way for agents to compare dozens or hundreds of low-cost services without requiring them to register for each one, add payment details, or generate API keys. An agent could test several APIs for only a few cents each, while the owner limits its total exposure. The same controls could allocate recurring AI inference budgets to employees and flag unusually fast spending for human review. Wallets will supply the buyer side of Cloudflare’s broader agentic commerce push. Its Monetization Gateway is intended to enable eligible Cloudflare customers to sell APIs and content via x402-compatible endpoints, while Wallets will add purchasing support to the company’s Agents SDK. Initial funding options will include onramps and offramps in supported geographies, with self-funding in stablecoins planned as an alternative for eligible users. Cloudflare also plans to link wallets to accounts through cloudflare\[.\]pay, giving agents optional, persistent, human-readable identities such as research.example.cloudflare.pay. The identifier would serve as a readable label for a cryptographic key pair, building on Cloudflare’s Web Bot Auth work. Merchants could decide whether to favor or require identified agents, while agents could choose whether to declare who they represent. Together, payments, spending controls, and optional identity are meant to give agents a machine-native route to buy services while keeping people in control of budgets and exceptions. [Source](https://blog.cloudflare.com/wallets/?ref=testingcatalog.com) ### Mistral releases Shieldstral for multimodal moderation URL: https://www.testingcatalog.com/mistral-releases-shieldstral-for-multimodal-moderation/ Last updated: 2026-08-05T11:53:26.000Z Mistral has released Shieldstral, a 3B open-weight multimodal safety classifier for teams that need moderation rules tailored to a product, audience, or domain. Released under Apache 2.0, the model handles text, images, and combined text-image content through one interface and can run on a single 16GB NVIDIA GPU. Its weights are available from Mistral on [Hugging Face](https://huggingface.co/mistralai/Shieldstral-1.0-3B?ref=testingcatalog.com). Shieldstral turns moderation into a binary question-answering task. At inference time, a developer supplies the evaluation context, a plain-language yes-or-no policy question, and the content to assess, which can be a prompt, a response, a prompt-response pair, or an image with optional text. The model reads the yes and no logits and converts them into a continuous safety score, allowing applications to set thresholds or rank results by confidence. Policies remain in the prompt, so a single checkpoint can be retargeted without retraining. > 🛡️Introducing Shieldstral, Mistral’s 3B open-weights model for content safety that can be deployed on-device 🧵 [https://t.co/SXBTWIJO9p](https://t.co/SXBTWIJO9p?ref=testingcatalog.com) [pic.twitter.com/Hhjw8o0gft](https://t.co/Hhjw8o0gft?ref=testingcatalog.com) > > — Mistral AI (@MistralAI) [August 4, 2026](https://x.com/MistralAI/status/2084684735725379637?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Mistral says Shieldstral matches or exceeds open-guard models up to 7 times larger in text safety, refusal detection, policy adaptability, and multimodal moderation. The company also reports state-of-the-art multimodal results, with all evaluation samples held out from training. The same setup supports prompt classification, response moderation, refusal checks, and toxicity detection. The model was trained on real and synthetic sources whose different labels and taxonomies were converted into a shared instruction-query-document format. Mistral used contrastive examples to teach distinctions between closely related policies, supplemented scarce visual safety data with general image datasets as negative examples, and filtered image-query pairs with a vision-language reranker. LoRA fine-tuning and SLERP then combined checkpoints focused on public-data calibration, fine-grained policy discrimination, and base-model instruction following. Mistral built Shieldstral on Forge, its platform for training, aligning, and evaluating custom models. The release also advances the company's work as an inaugural member of the Open Secure AI Alliance alongside NVIDIA and other organizations. Mistral plans to extend Shieldstral with broader multilingual coverage, stronger support for long documents, and broader multimodal safety capabilities. [Source](https://mistral.ai/news/shieldstral/?ref=testingcatalog.com) ### Alibaba released Qwen3.8-Max with open weights coming soon URL: https://www.testingcatalog.com/qwen-released-qwen3-8-max-with-open-weights-coming-soon/ Last updated: 2026-08-04T22:21:30.000Z Qwen has officially released Qwen3.8-Max, its most capable model to date and its first at Max scale slated for open weights. The model is available now through QwenCloud, while its weights are due on Hugging Face and ModelScope next week. Built on the Qwen3.5 architecture, it has 2.4 trillion parameters with 95 billion active and is aimed at coding, research, professional work, and multimodal tasks. > 📢Meet Qwen3.8-Max — our most capable model to date. > > Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!🎉 > > Qwen3.8-Max, a new bar for coding and cowork at 2.4T parameters: > > \- Autonomous coding: 10+ days of… [pic.twitter.com/e3YFj2hqcT](https://t.co/e3YFj2hqcT?ref=testingcatalog.com) > > — Qwen (@Alibaba\_Qwen) [August 3, 2026](https://x.com/Alibaba%5FQwen/status/2084100707423289643?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Autonomous coding sits at the center of the launch. In one company-run test, Qwen3.8-Max spent about 16 days building and maintaining the oh-my-cli project, producing 265 commits, 127 pull requests, and 151 issues through a loop of task intake, implementation, testing, and repair. In another, it recreated a research pipeline, completed 33 GPU training rounds, and developed a method that Qwen says scored 2.7 points above the paper's approach on AIME24. Qwen is also pitching the model for long-running professional workflows. During an autonomous hardware-design test, it reduced a cryptographic accelerator from 8,298 to 678 gates over roughly 500 turns and reached timing closure at 500 MHz after physical layout. In a year-long ecommerce simulation spanning more than 2,000 rounds, it finished with ¥416,252, a 4.16-times return and 38% more than the runner-up. > Cowork: Any role, build beyond [pic.twitter.com/Hggkj3DFSs](https://t.co/Hggkj3DFSs?ref=testingcatalog.com) > > — Qwen (@Alibaba\_Qwen) [August 3, 2026](https://x.com/Alibaba%5FQwen/status/2084100723558797458?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The multimodal system is designed to process documents longer than 200 pages and videos exceeding 100 hours, organize their contents into traceable structures, and visually inspect its own output for corrections. Qwen is also introducing Qwen-MM-Plugins, a harness extension for multimodal memory and visual tool use in areas including video editing, Blender, and CAD. RecreationBench combines code and GUI operation to rebuild black-box applications across desktop, mobile, and web platforms. The Qwen team is pairing the release with familiar deployment routes for developers, researchers, and enterprises. QwenCloud supports OpenAI-compatible chat completions and Responses APIs alongside an Anthropic-compatible interface, allowing use with popular agent frameworks and coding assistants. API users can select xhigh, medium, or low reasoning effort, with thinking preserved by default. Releasing the weights next week will bring a Max-scale Qwen model into self-hosted and research settings for the first time. [Source](https://qwen.ai/blog?id=qwen3.8&ref=testingcatalog.com) ### Google is working on Plugins for Gemini Enterprise URL: https://www.testingcatalog.com/google-is-working-on-plugins-for-gemini-enterprise/ Last updated: 2026-08-03T13:56:00.000Z Google appears to be preparing Plugins and a dedicated Notifications area for the Gemini Enterprise app, extending its push beyond chat toward reusable workplace workflows. Clues in an unfinished interface show a revised Connectors area split into three tabs: Connectors, Skills, and Plugins. The Plugins tab is currently empty, indicating that the feature is still under development. Its placement suggests that Plugins may act as mini-apps or packaged workflows, combining reusable Skills with one or more Connectors. This could let business teams run multi-step processes across company data and services without configuring every component from scratch. The structure also resembles how [Claude](https://www.testingcatalog.com/anthropic-adds-plugins-support-for-claude-cowork-on-paid-plans/) separates connectors and skills, while allowing packaged tools to build on both. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Gemini-Enterprise-08-01-2026_11_52_PM.jpg) A separate Notifications area is also being prepared. It is expected to collect completed Gemini responses from scheduled tasks and conversations placed in a queue, giving users one place to return to work that finished in the background. This would be particularly useful for employees running research, reporting, or other jobs that take longer than a standard chat response. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Gemini-Enterprise-08-01-2026_11_51_PM.jpg) The work fits Google’s broader strategy for [Gemini Enterprise](https://www.testingcatalog.com/tag/gemini/), which brings agents, company data, and workflow controls into one governed service. Google already describes Skills as reusable actions, Connectors as links to services including Google Workspace and Microsoft 365, and Inbox as a central hub for long-running agent activity. Plugins would add a packaged application layer on top of those pieces. 💡 Join [TestingCatalog AI News](https://t.me/testingcatalog?ref=testingcatalog.com) on Telegram! Google has not publicly announced the Plugins section or provided a rollout date. The empty catalog and unfinished Notifications interface indicate that both are still in development, with availability, supported accounts, and the first plugin partners remaining unknown. ### Exclusive: Microsoft tests new MAI Realtime voice model URL: https://www.testingcatalog.com/exclusive-microsoft-tests-new-mai-realtime-voice-model/ Last updated: 2026-08-26T18:51:53.000Z Microsoft appears to be preparing its first native real-time voice model, referred to as MAI Realtime, which has surfaced as a hidden early-access entry in the company’s MAI Playground. The listing suggests a small group of partners already has hands-on access, and what is visible points to a bidirectional, full-duplex system, one that listens and speaks at the same time rather than trading turns, placing it in direct comparison with [OpenAI’s GPT Live 1](https://www.testingcatalog.com/openai-rolls-out-gpt-live-voice-for-chatgpt-on-web-and-mobile/) or [Sesame](https://www.testingcatalog.com/sesame-debutes-ios-app-in-preview-with-personal-voice-agents/). > \> It supports English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese, and Thai. [pic.twitter.com/7BmcW2MBa0](https://t.co/7BmcW2MBa0?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [August 2, 2026](https://x.com/testingcatalog/status/2083928447450067196?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Two voices are present so far, Victoria and Grant, both noticeably more natural than what Copilot’s voice mode currently delivers. Language can be pinned explicitly or left on automatic detection, and the model switches languages mid-conversation without losing its footing. Turn-taking is configurable through two listener options: a **Switchboard mode** built around an MAI-Ears endpointer driven by inline control tokens, and a **deterministic setup** that pairs silence-based endpointing with a Whisper semantic endpointer. ![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_34_PM.jpg) ![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_33_PM--1-.jpg) The practical difference between them is subtle in use, though interruptions are handled cleanly and response latency is low. The model does not sing or produce non-speech sounds, which keeps it squarely a conversational system rather than a general audio generator. A debug panel exposes live latency figures, model thoughts and processing steps, and sample sharing looks set to arrive for playground users once access widens. ![MAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/MAI-Playground-Microsoft-AI-08-02-2026_04_33_PM.jpg) That would fill a conspicuous gap. Every [MAI speech model](https://www.testingcatalog.com/microsoft-previews-mai-image-2-5-pro-and-mai-voice-2-flash/) shipped so far runs in one direction: MAI-Voice-2 and its Flash variant for synthesis, and MAI-Transcribe-1.5 for recognition, while the speech-to-speech layer in Azure Speech’s Voice Live API still relies on the GPT-Realtime model. A first-party full-duplex model would close that dependency for Mustafa Suleyman’s superintelligence team, which shipped seven in-house models at Build 2026 and has been steadily swapping OpenAI components out of Copilot, Teams and Bing. [Microsoft](https://www.testingcatalog.com/tag/microsoft/) Foundry is the likely developer destination, with Copilot voice the obvious consumer surface, though no timeline has been attached to either. ### Google is aiming to close feature gaps on Gemini desktop URL: https://www.testingcatalog.com/google-is-aiming-to-close-feature-gaps-on-gemini-desktop/ Last updated: 2026-08-01T22:49:24.000Z Google is narrowing the gap between the Gemini desktop app and its web counterpart, with several unreleased changes now in testing. Dedicated tabs for **image and video generation** are being added to the app's navigation, and an app configuration panel has appeared on the desktop for the first time. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-07-29-at-16.11.04.webp) A second track is a **camera** attachment. From the attachment menu, users would select a camera option, which opens a capture screen, takes a picture, and drops it into the prompt. It reads as photo capture rather than a live video stream, keeping it distinct from [Gemini](https://www.testingcatalog.com/tag/gemini/) Live's real-time sharing and closer to what mobile users already have. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/08/Screenshot-2026-07-29-at-16.16.49.webp) The third piece sits with [Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/). Its connectors menu on desktop now includes an option to add custom MCP servers, bringing something previously confined to internal configuration into the visible agent settings. Google already permits custom MCP apps for Spark via the web app, so this is an extension of that rather than the first appearance, though it matters for anyone running Spark from the Mac app. > Gemini Spark is rolling out to Google AI Pro users outside the U.S. > > Spark is your personal AI agent that works in the background 24/7 to get things done under your direction, handling the heavy lifting so you can focus on what matters. [https://t.co/qq1RQK221m](https://t.co/qq1RQK221m?ref=testingcatalog.com) > > — Google Gemini (@GeminiApp) [July 31, 2026](https://x.com/GeminiApp/status/2083302569796059271?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Timing is unclear; these are trusted-tester builds with no public rollout date. The context is a busy month: Spark reached macOS on July 1 alongside Google Tasks and Keep connections, and intelligent dictation and screen-aware reasoning began rolling out on July 29\. Together, generation tabs, camera capture, and connector management would leave fewer reasons to open the browser. ### xAI adds character references and 1080p to Imagine Video 1.5 URL: https://www.testingcatalog.com/xai-adds-character-references-and-1080p-to-imagine-video-1-5/ Last updated: 2026-08-01T09:58:28.000Z xAI has expanded Imagine Video 1.5 with image and voice references, prompt-only video generation, and native 1080p output. The update gives users several ways to shape a clip: describe a shot without a starting image, guide it with reference images, or combine a character image with a voice sample. Image and voice references are starting in the United States for SuperGrok Heavy and SuperGrok Plus subscribers on Grok’s Imagine website and iOS app. xAI says the tools will roll out to all tiers over the next few days. Text-to-video and native 1080p generation are already generally available through Grok Imagine on the web, iOS, and Android, with 1080p supported for both text-to-video and image-to-video workflows. > Imagine Video 1.5 launched last month as our best video model yet, with more lifelike motion and sound. Today it goes even further with text-to-video support, image and voice references, and native 1080p.[https://t.co/OWP5rCxlru](https://t.co/OWP5rCxlru?ref=testingcatalog.com) [pic.twitter.com/bHUf5pz62S](https://t.co/bHUf5pz62S?ref=testingcatalog.com) > > — Grok (@grok) [August 1, 2026](https://x.com/grok/status/2083353607370416632?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Multi-Reference lets creators assign separate visual anchors to a generation. One image can hold a face in place, another can preserve a product, and another can define a location. Users can keep a character while changing the scene, retain a scene while swapping the character, or preserve both while altering only the action. The model accepts up to seven references in a single generation. 0:00 /0:11 1× Voice consistency adds another layer of continuity. When a character image and voice reference are supplied together, xAI says the model can maintain the same face and voice across scenes. That targets a common production need: generating multiple shots around a recurring character without rebuilding identity cues for every clip. Developers can access image references, text-to-video, and native 1080p through the xAI API using the grok-imagine-video-1.5 model. Voice references require a request to xAI. The API example shows a prompt, reference image URL, duration, aspect ratio, and resolution passed into a video generation call. The rollout extends the Imagine Video 1.5 release introduced last month, which xAI presented as an advance in motion, physics, and audio. The latest additions shift the focus toward tighter control over identity, scene, and product while expanding the model from image-led animation to videos generated directly from text. [Source](https://x.ai/news/grok-imagine-video-1-5-references?ref=testingcatalog.com) ### Thinking Machines launched open-weight Inkling-Small URL: https://www.testingcatalog.com/thinking-machines-launched-open-weight-inkling-small/ Last updated: 2026-08-01T09:32:34.000Z Thinking Machines Lab is releasing Inkling-Small, an efficient, open-weight model designed to deliver performance comparable to Inkling while being one-quarter of its size. The Mixture-of-Experts transformer contains 276 billion total parameters, with 12 billion active. Its full weights are being released on Hugging Face, with fine-tuning offered through Tinker and text, image, and audio chat available in Tinker Playground. The model combines a context window of up to one million tokens with variable thinking effort from minimal to xhigh, allowing developers to trade compute for performance. It was trained on NVIDIA GB300 NVL72 systems and uses substantially less compute than Inkling, according to Thinking Machines. That lower requirement positions it for experimentation, coding, tool use, fine-tuning, and real applications. > Today, we are releasing Inkling-Small. > > Inkling-Small achieves comparable performance to Inkling at a quarter of its size. It features 276B total parameters, 12B active. We are making the full weights available.[https://t.co/NzFVYVkuQI](https://t.co/NzFVYVkuQI?ref=testingcatalog.com) > > Fine-tune it on Tinker today, or chat with… > > — Thinking Machines (@thinkymachines) [July 30, 2026](https://x.com/thinkymachines/status/2082885869426631032?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Thinking Machines began Inkling-Small after training Inkling, giving the lab room to revise its pre-training data mix and machine-learning recipe. An earlier preview checkpoint was post-trained partly through on-policy distillation with Inkling as the teacher, followed by two weeks of scaled agentic coding reinforcement learning. The lab says the resulting model overtakes Inkling on reasoning and agentic coding benchmarks, although Inkling remains stronger in knowledge coverage and factuality. Reported results include 31.6 percent on the text-only Humanity's Last Exam, compared with 29.7 percent for Inkling. Inkling-Small also reached 80.2 percent on SWEBench Verified, 55.9 percent on the public SWEBench Pro evaluation, and 64.7 percent on Terminal-Bench 2.1 with its best harness. Thinking Machines ran its evaluations at an effort of 0.99 and a temperature of 1.0, with disclosed caveats regarding internal harness and formatting. Inkling-Small keeps Inkling's natively multimodal, encoder-free design. Audio is converted to dMel spectrograms, while images are split into 40-by-40-pixel patches and transformed by a four-layer hMLP before being processed with text tokens. The model can use Python for visual work, combining reasoning with cropping, zooming, and programmatic inspection of dense documents and charts. Safety post-training follows Inkling's recipe and is backed by internal testing and red-teaming by trusted external partners. Thinking Machines reports 98.4 percent on StrongREJECT, a 71.6 percent adversarial refusal rate on FORTRESS, and a 96.9 percent benign answer rate. Inkling and Inkling-Small are currently offered on Tinker with a limited-time discount. [Source](https://thinkingmachines.ai/news/inkling-small/?ref=testingcatalog.com) ### Gemini for macOS adds new "Speak to Window" feature URL: https://www.testingcatalog.com/gemini-for-macos-adds-new-speak-to-window-feature/ Last updated: 2026-07-31T21:32:28.000Z Google has begun rolling out a [new voice capability](https://www.testingcatalog.com/google-tests-voice-dictation-and-magic-pointer-on-gemini-desktop/) for Gemini on macOS, allowing people to create, edit and summarize content while staying inside the app or desktop window already in use. The launch is global for all Gemini app users on Mac in English, with more languages due later. A long-press of the Fn key opens voice input in any desktop window and places the result at the cursor. > Simply ramble into the Gemini macOS app and let Gemini's powerful audio understanding capabilities get to work. > > \- Extract and summarize information > \- Compose and rewrite text > \- Generate and edit images > > Voice UI is the future and I'm here for it 🙌 > > Long-press Fn key and 💬 [pic.twitter.com/IcLoZp9Qvp](https://t.co/IcLoZp9Qvp?ref=testingcatalog.com) > > — Thor 雷神 ⚡️ (@thorwebdev) [July 29, 2026](https://x.com/thorwebdev/status/2082527440279130375?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The default intelligent-dictation mode does more than produce a literal transcript. It turns speech into polished text, removes filler words, respects corrections made during a sentence and formats the final output. Because the text appears where the user is already working, drafting and revision can remain in the current window rather than requiring a switch to a separate Gemini conversation. A separate option in [Gemini](https://www.testingcatalog.com/tag/gemini/) settings enables reasoning. In this mode, Gemini uses visible on-screen context for more involved voice requests. A user can highlight local files, images or documents and ask for extraction or a summary, select text and request a rewrite or tone change, or create and edit images through spoken instructions that refer to content on the desktop. The announcement comes from Michael Friedman, group product manager for the Gemini app, and Alvin Zhou, senior product manager at Google DeepMind. The Gemini app is available for macOS as the rollout reaches English-language users worldwide. [Source](https://blog.google/innovation-and-ai/products/gemini-app/speak-naturally-gemini-app-mac-os/?ref=testingcatalog.com) ### SpaceXAI launches Grok Voice Think Fast 2.0 on Agent Builder URL: https://www.testingcatalog.com/spacexai-launches-grok-voice-think-fast-2-0-on-agent-builder/ Last updated: 2026-07-29T21:11:15.000Z xAI has introduced Grok Voice Think Fast 2.0, its next-generation speech-to-speech model, with gains in intelligence, transcription accuracy, conversational behavior, and tool use. The model is aimed at developers [building voice agents](https://www.testingcatalog.com/icymi-xai-debuts-grok-voice-agent-builder-for-enterprises/) and costs $0.08 per minute of audio. xAI expects it to raise performance across almost all use cases without changes to existing prompts. On Artificial Analysis’ speech-to-speech benchmark, Think Fast 2.0 scored 82.9% overall, up from 75.7% for version 1.0 and ahead of GPT-Realtime-2.1 at 79.1% and Gemini 3.1 Flash at 69.5%. Its agentic score reached 56.5%, compared with 52.1% for its predecessor, 45.7% for GPT-Realtime-2.1, and 37.7% for Gemini 3.1 Flash. Time to first audio fell from 1.25 seconds to 0.70 seconds. The model’s 95.1% conversational benchmark score was just below GPT-Realtime-2.1 at 95.7%. > Announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.[https://t.co/XUiX1CouKz](https://t.co/XUiX1CouKz?ref=testingcatalog.com) [pic.twitter.com/Nel3zwzkwY](https://t.co/Nel3zwzkwY?ref=testingcatalog.com) > > — SpaceXAI (@SpaceXAI) [July 29, 2026](https://x.com/SpaceXAI/status/2082529280341553209?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Transcription is another major focus. In xAI’s evaluation of thousands of short phrases across 24 languages, the company reported accuracy improvements of 1.5 to 2.0 times versus Deepgram Nova 3 and ElevenLabs Scribe v2, and 1.4 times versus Think Fast 1.0\. xAI says the gap is roughly 10× under substantial background noise and telephony compression. The comparison uses the word error rate, with lower scores preferred. Think Fast 2.0 reasons in parallel with speech, a design intended to preserve latency while handling more complex queries. Median relative reasoning-token use fell to 0.4 times, using the predecessor’s 1.0 times as a baseline. xAI says this lets production tool calls usually execute before the agent finishes its first sentence. Reinforcement learning also pushed the model toward shorter sentences, one question at a time, and less fluff while guiding users through complex workflows. For [SpaceXAI](https://www.testingcatalog.com/tag/grok/), this release is a push to make Grok Voice more dependable in real customer workflows. An A/B test on Starlink’s phone service produced higher sales conversion and support containment rates, according to the company. On August 5, 2026, the grok-voice-latest alias will automatically move from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0\. Developers who want the prior model must pin the 1.0 identifier before the switch; everyone else needs to take no action. [Source](https://x.ai/news/grok-voice-think-fast-2?ref=testingcatalog.com) ### Vellum adds Voice Mode to personal AI assistant URL: https://www.testingcatalog.com/vellum-adds-voice-mode-to-personal-ai-assistant/ Last updated: 2026-07-29T19:43:47.000Z Vellum has introduced Voice Mode for its personal assistant, enabling spoken conversation to be processed through the same agent loop that powers text chat. Speech is transcribed using Deepgram, processed by the assistant with full access to its tools, memory, and skills, and then spoken back through ElevenLabs. The assistant can browse the web, read files, run code, send messages, and manage a calendar while the conversation is ongoing, maintaining the same identity across both text and voice interfaces. > Work will look like this sooner than you think. > > You can now speak to your [@vellum\_ai](https://x.com/vellum%5Fai?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) agent to handle complex tasks like plan your trip, push a PR, or optimize your Meta ads. > > Watch Alex get things done with just his voice: [pic.twitter.com/28vCTVWo8Z](https://t.co/28vCTVWo8Z?ref=testingcatalog.com) > > — anita · vellum.ai 👾🦾 (@anitakirkovska) [July 29, 2026](https://x.com/anitakirkovska/status/2082524426717782222?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Vellum opted to retain the classic cascade architecture instead of transitioning to a native speech-to-speech model. This decision was based on the belief that a closed audio-in, audio-out model does not allow for an agent loop, tools, memory, or approvals. The engineering effort focused on seamlessly integrating these components. A speculative launch mechanism initiates the agent's response as soon as the voice detector identifies trailing silence, running the endpoint decision concurrently with the time to first token, and rolling back the turn if the speaker was merely pausing mid-thought. Each turn first engages a fast front model that either responds directly, remains silent if the speaker seems unfinished, or escalates to a more robust model while delivering a brief holding phrase audibly. During lengthy tool calls, the assistant narrates progress in audio only, ensuring that this narration does not persist in the transcript. ![Vellum](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Screenshot-2026-07-29-at-21.31.46.webp) Interruption handling is a feature Vellum describes as custom engineering rather than an API flag. Interruptions such as cutting in mid-sentence, redirecting, or adding information while the assistant is processing are managed effectively, with the assistant yielding and maintaining context. A guard mechanism waits for a quarter second of sustained speech before yielding, preventing the assistant's audio from being misinterpreted as an interruption due to browser echo cancellation. Synthesis operates seamlessly, with the opening clause spoken as soon as it is ready and a prefetch pump ensuring the next segment is ready behind the current playback. 0:00 /0:46 1× Voices are managed through a Vellum Managed Voice option that automatically handles ElevenLabs voices, with the option to provide an ElevenLabs key for custom voices. If no key is provided, the system defaults to managed voices. Voice continuity is maintained across devices, allowing a task started on a phone to continue on a Mac with the same assistant and memory. Complex multi-step processes automatically transition to text chat when text becomes the more suitable medium. SPONSORED Start testing Vellum AI [Learn more ](https://www.vellum.ai/product?ref=testingcatalog.com) Vellum's personal assistant is designed with its own identity, memory, and skills, operating on macOS, iOS, web browsers, Slack, Telegram, and more! There are also plans for future Android and Windows support. The assistant comes equipped with over 60 skills, a permission system ranging from strict to full access, and credentials stored in the macOS Keychain when self-hosted or in an isolated vault on the managed platform. The core assistant is open source, and the company is led by Chief Executive Akash Sharma, alongside Co-Chief Technology Officers Noa Flaherty and Sidd Seethepalli. ### Bagel Labs launches WorldDiT world model for robotics URL: https://www.testingcatalog.com/bagel-labs-launches-worlddit-world-model-for-robotics/ Last updated: 2026-07-29T13:16:46.000Z Bagel Labs has released WorldDiT, a world model for robotics that learns two things at once: the action a robot should take, and how the scene in front of it is likely to change next. Both abilities use the same shared parameters rather than being split across separate models. The release targets robotics engineers, researchers, and technical leads at robotics companies and includes a technical report and model weights published on [Hugging Face](https://huggingface.co/bageldotcom/worlddit?ref=testingcatalog.com). The architecture question behind it is a live one in robot learning. Many current systems either train a policy that predicts actions alone, or lean on a very large vision-language backbone to hold perception, language, and control together in a single stack. WorldDiT takes the unified route at a much smaller scale, treating prediction and control as two outputs of one trained system, so the parameters that model how a scene evolves are the same parameters that decide what the robot does about it. > WorldDiT learns to sample robot actions and predict a future view of the world in parallel through one diffusion backbone. That richer signal helps it reach strong LIBERO results with far fewer parameters. > > Table below shows the full benchmark comparison. [pic.twitter.com/ML9De2qkQL](https://t.co/ML9De2qkQL?ref=testingcatalog.com) > > — bagel.com (@bageldotcom) [July 28, 2026](https://x.com/bageldotcom/status/2082179137813586404?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The number Bagel Labs is putting forward is size. On LIBERO, a standard simulation benchmark for language-conditioned robot manipulation, the company reports that WorldDiT sits on the frontier of benchmark score relative to parameter count while staying well below one billion parameters. Framed that way, the claim is less about topping a leaderboard outright and more about the ratio behind it: strong robot simulation results without scaling into multi-billion parameter territory. For teams running inference on the robot instead of in the cloud, that ratio carries weight, since memory budget and latency on embedded hardware both sit downstream of model size. Bagel Labs is an AI research lab working on distributed training of frontier diffusion models across heterogeneous and commodity hardware. Its previous releases follow the same thread. Paris, published in October 2025, was an open-weight decentralized diffusion model for image generation, built from expert models trained in isolation with no gradient, parameter, or activation exchange between them, then combined at inference by a lightweight router. Paris 2.0 followed in May 2026 and carried that recipe into video, with the lab reporting Frechet Video Distance falling from 561.04 to 279.01 against a monolithic model trained on matched data and compute. A related paper on heterogeneous decentralized diffusion models was accepted at CVPR 2026. WorldDiT moves that work from media generation into physical AI, a direction the lab has been signaling since Paris 2.0, when it noted that video diffusion backbones increasingly sit underneath the world models used in robotics. Bagel Labs was founded by Bidhan Roy and publishes its research and model weights in the open, with the WorldDiT technical report released alongside the model. [Source](https://blog.bagel.com/p/worlddit?ref=testingcatalog.com) ### Google is working on interactive Apps for Gemini Notebook URL: https://www.testingcatalog.com/google-is-working-on-interactive-apps-for-gemini-notebook/ Last updated: 2026-07-29T12:08:04.000Z Google is preparing a new artifact type for Gemini Notebook that would turn a set of sources into a working interactive app. A tile marked "App" has surfaced in the Studio panel next to the still-unreleased [Canvas](https://www.testingcatalog.com/google-tests-canvas-and-connectors-on-notebooklm/) and [Lit Review](https://www.testingcatalog.com/google-tests-literature-review-matrix-tool-for-notebooklm/) options, described as generating interactive applications from a notebook's material. It accepts a prompt, so the output could be a dashboard, a study aid, or a lightweight game built from a book. Nothing here is live for anyone yet, and no release window has appeared. The timing lines up with what Google shipped this month. On 16 July, it renamed NotebookLM to Gemini Notebook and extended the per-notebook secure cloud computer to AI Pro on the web, after an Ultra-first rollout in June. That sandbox writes and runs code against uploaded sources, the substrate any generated app would need to execute at all. ![Gemini Notebook](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Gemini-Notebook-07-29-2026_12_17_AM.jpg) Smaller changes sit alongside it. [AI Notes](https://www.testingcatalog.com/google-prepares-personalization-and-ai-editing-for-notebooklm/) keep getting refined, with a manual add-note button joining the previously spotted option to have the model draft one. A watermarking toggle has also turned up, which may govern the visible Google branding stamped on exported videos, slide decks and infographics, though its scope stays unclear. ![Gemini Notebook](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/TestingCatalog-Latest-AI-Model-Releases-and-Agent-News-Gemini-Notebook-07-29-2026_10_56_AM.jpg) Notice labels point to a model upgrade for the notebook and to web sourcing pulled mid-chat, with those results savable as sources afterward. Which model is not stated; the flagship Pro tier has slipped repeatedly, while a newer Flash release landed on 21 July. ![Gemini Notebook](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/TestingCatalog-Latest-AI-Model-Releases-and-Agent-News-Gemini-Notebook-07-29-2026_12_25_AM.jpg) Students, academics, and analysts are the obvious audience. Google has spent a year turning [Gemini Notebook](https://www.testingcatalog.com/tag/notebooklm/) into an output factory: audio and video overviews, infographics, slide decks, flashcards, data tables, and an App tile would give it the first artifact a reader can operate rather than only consume. ### Cursor launches Start plan in India for ₹649 per month URL: https://www.testingcatalog.com/cursor-launches-start-plan-in-india-for-649-per-month/ Last updated: 2026-07-29T12:07:31.000Z Cursor has launched Cursor Start, a new individual plan built for developers in India. Available now for ₹649 per month, tax included, it brings expanded access to Grok 4.5 and Composer alongside more agent requests than the Free plan. Billing is monthly in Indian rupees, and customers can pay by UPI, credit card, or debit card. The plan covers Cursor across desktop, web, iOS, and the command line. Subscribers can use Grok 4.5, which Cursor describes as its most powerful model, and Composer, its most price-efficient coding model. Start also includes always-on cloud agents that can run longer jobs, build and test software, and open pull requests while the developer continues other work. From Cursor for iOS, users can launch agents or control existing ones before returning to the same work on desktop. Plugins, MCP servers, hooks, and skills can extend the setup across a broader development workflow. > Today we're launching Cursor Start, a new ₹649/month plan for developers in India. > > Start includes generous access to Grok 4.5 and Composer, so you can plan, build, test, and ship with agents every day. [pic.twitter.com/y2ZYlyzfR6](https://t.co/y2ZYlyzfR6?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [July 28, 2026](https://x.com/cursor%5Fai/status/2081978255004053560?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Start sits between Cursor's Free and Pro tiers. Free provides Composer and a limited monthly allowance of local agent requests without payment. Pro adds every major model, including advanced systems from other labs, as well as Bugbot, Auto mode, Automations, the Cursor SDK, and on-demand usage beyond included limits. Start instead targets routine daily building with a lower local price and a narrower model selection. The India focus follows rapid growth for Cursor in the country. Cursor says its Indian user base tripled over the past year to more than three million developers, making India its third-largest market. It also reports that India has more power users than any other market and the highest number of agent requests per developer. The company positions Start as a response to repeated requests for local-market pricing and UPI support from students, founders, engineers, and designers. Existing Free users in India can upgrade through the Cursor dashboard. New users can select Start during signup. The move gives Cursor a region-specific tier for a market already driving unusually heavy agent use, while leaving Pro as the step up for developers who need a wider model catalog and additional automation tools. [Source](https://cursor.com/blog/cursor-start-india?ref=testingcatalog.com) ### Fish Audio launches S2.1 Pro with support for 83 languages URL: https://www.testingcatalog.com/fish-audio-launches-s2-1-pro-with-support-for-83-languages/ Last updated: 2026-07-28T17:29:21.000Z Fish Audio has positioned S2.1 Pro as its recommended production voice model, built around real-time conversational speech rather than the scripted narration that older text-to-speech systems were designed for. The model is available now through the Fish Audio API, alongside a free tier that runs the same model at no cost for development and testing under fair-use limits. The company presents it as its most capable production model so far, timed to a company anniversary. The technical pitch centers on latency and language coverage. S2.1 Pro reports a time-to-first-audio of around 90 milliseconds on standard calls, fast enough for natural turn-taking in live dialogue, and it covers 83 languages with a single voice identity that holds across them. Delivery is steered through free-form bracket tags written directly into the text, so a line can carry instructions like a whisper or a nervous laugh without switching to a fixed menu of preset emotions. The model also supports multi-speaker dialogue and voice cloning from short reference samples, typically in the range of 10 to 30 seconds, capturing tone and speaking style without additional fine-tuning. ![Fish Audio](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Best-AI-Text-To-Speech-Free-Voice-Cloning-Fish-Audio-07-22-2026_11_41_AM--1-.jpg) S2.1 Pro is built as an improvement over the earlier S2-Pro across quality, latency, and throughput, and it carries production guarantees around time-to-first-audio and data processing for teams running it at scale. It plugs into agent workflows through MCP and agent-skill support, which points the model at developers wiring speech into voice agents, phone systems, and long-form audio pipelines. The free tier lowers the barrier for smaller teams and prototyping, offering the same model quality and language coverage as the paid production string without a hard usage cap. SPONSORED Start testing Fish Audio [Learn more ](https://fish.audio/?ref=testingcatalog.com) Fish Audio builds speech models for creators, developers, and enterprises, with a platform spanning voice generation, voice cloning, and real-time voice applications. Its open-source Fish Speech line has drawn a developer following, passing 20,000 stars on GitHub, and the company has since moved through the OpenAudio S1 generation into the S2 family, which shipped with open weights earlier in the year. The S2 models were trained on large multilingual audio datasets and, by the company's own benchmarks, reached low word error rates against other evaluated systems. S2.1 Pro is the production layer on top of that research, aimed at the growing set of products where voice is becoming the primary interface and latency, language reach, and cloning fidelity decide whether an assistant feels present or delayed. ### Anthropic launched Claude Opus 5 across all platforms URL: https://www.testingcatalog.com/anthropic-launched-claude-opus-5-across-all-platforms/ Last updated: 2026-07-26T07:57:01.000Z Anthropic has launched Claude Opus 5, now available across all of its platforms. It is the new default model for Claude Max and the strongest model on Claude Pro. Developers can access claude-opus-5 through the Claude API. Pricing remains at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8, while fast mode runs at about 2.5 times the default speed for twice the base price. > Introducing Claude Opus 5. > > It's a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price. [pic.twitter.com/GQWhcq2CQL](https://t.co/GQWhcq2CQL?ref=testingcatalog.com) > > — Claude (@claudeai) [July 24, 2026](https://x.com/claudeai/status/2080699495453528290?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Built for daily coding and knowledge work, Opus 5 is designed to approach Claude Fable 5's intelligence at half the price. Its effort setting lets customers trade token use and cost against speed and capability. Anthropic says the model leads Frontier-Bench v0.1 and more than doubles Opus 4.8's result at a lower cost per task. At maximum effort on CursorBench 3.2, it finished within 0.5% of Fable 5's peak score for half the cost per task. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/7530b1086992936d7e9d5796a892d1e8fa063253-3840x2160.webp) Anthropic also reports that Opus 5 scored three times as high as the next-best model on ARC-AGI 3 and beat every other model at any given cost on OSWorld 2.0\. It surpassed Opus 4.8 across Anthropic's life-sciences evaluations and produced stronger visual output. The model places more emphasis on checking its work: in one example, it built a computer-vision pipeline to reconstruct a machine-part drawing as a 3D FreeCAD model after repeated attempts. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/b5e071ba6a9ce5628b4662f05484d1806a9fdc94-3840x2160.webp) Early-access customers reported strengths in difficult debugging, analytics, financial research, genomics, legal agent work, code review, and long-running coding tasks. Anthropic says Opus 5 is its most aligned model so far and does not move the frontier in risky dual-use capabilities. It remains behind Mythos 5 in biology research and offensive cybersecurity. For Anthropic, the release pairs wider access with narrower cyber controls. Classifiers permit vulnerability discovery in source code but block binary scanning, penetration testing, and exploit generation; flagged requests fall back to Opus 4.8 by default in Claude AI, Claude Code, and Claude Cowork. Two beta updates also let API developers change tools mid-conversation without invalidating the prompt cache and route safety-flagged requests to another model. [Source](https://www.anthropic.com/news/claude-opus-5?ref=testingcatalog.com) ### ICYMI: Black Forest Labs opens FLUX 3 Video early access URL: https://www.testingcatalog.com/icymi-black-forest-labs-opens-flux-3-video-early-access/ Last updated: 2026-07-26T07:40:27.000Z Black Forest Labs has opened early access to FLUX 3 Video, the first available part of a new multimodal foundation model trained jointly on images, video, and audio. The model can create videos with native audio up to 20 seconds long in one generation, starting from text, images, keyframes, or reference clips. It can also continue video and audio, carry central elements such as a character into new scenes, produce multilingual dialogue, and chain clips into longer multi-shot sequences. > Introducing FLUX 3. > > One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. > > FLUX 3 Video is now available in early access (link below). > > Jointly trained in one unified architecture, our model can be extended to… [pic.twitter.com/voQ5iUJJZY](https://t.co/voQ5iUJJZY?ref=testingcatalog.com) > > — Black Forest Labs (@bfl\_ai) [July 23, 2026](https://x.com/bfl%5Fai/status/2080308988961554582?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Rather than treating each medium as a separate task, FLUX 3 uses one architecture to learn how appearance, motion, and sound constrain the same event. Language connects those perceptions to instructions and goals. The system builds on Black Forest Labs' Self-Flow method for aligning multimodal generation and understanding, with training scaled across all three media types at once. In preliminary tests, Black Forest Labs generated 10-second, 720p text-to-video clips with audio. The company says FLUX 3 was preferred over Grok Imagine Video in up to 69% of comparisons, Kling v3 Pro in 60%, Seedance 2.0 and Gemini Omni Flash in 52%, Runway Gen-4.5 in 77%, and Luma Ray 3.2 in 93%. The results remain early because both the model and evaluation harness are still in development. 💡 [Early Access request form](https://tally.so/r/44d9NX?ref=testingcatalog.com) FLUX 3 also supports image synthesis and editing across varied styles, aspect ratios, and resolutions, with stronger handling of complex prompts and multilingual text than earlier FLUX versions in midtraining evaluations. Early access to FLUX 3 Image is due in the following weeks. Future access is planned through APIs and private weights, alongside an open-weight FLUX 3 Dev backbone for image, video, audio, and action prediction. Black Forest Labs is positioning FLUX 3 as a step toward models that can perceive, predict, and act across physical and digital environments. Its action work follows two paths: native action prediction inside FLUX 3, and specialist models finetuned from its video backbone with limited task-specific data. The first partner project, FLUX-mimic, combines that backbone with mimic robotics' robot-learning work and is being tested on production tasks at Audi. The longer-term goal is to bring perception, action, and language prediction into one model. [Source](https://bfl.ai/blog/flux-3?ref=testingcatalog.com) ### Microsoft previews MAI-Image-2.5-Pro and MAI-Voice-2-Flash URL: https://www.testingcatalog.com/microsoft-previews-mai-image-2-5-pro-and-mai-voice-2-flash/ Last updated: 2026-07-26T07:29:03.000Z Microsoft has opened public previews of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, expanding its in-house AI model family with one option tuned for maximum visual fidelity and another designed for high-volume speech. The models are available through Microsoft Foundry, with trials also offered in the MAI Playground, and are aimed at creative teams, developers and enterprise contact centers. MAI-Image-2.5-Pro is Microsoft’s highest-fidelity image model to date. It is built for hero artwork, detailed editing and precise text inside generated images, with natural-language commands supporting rapid revisions. Pricing is $5 per 1 million text input tokens, $8 per 1 million image input tokens and $106 per 1 million image output tokens. > MAI-Voice-2-Flash launches today! Flash is 2x faster than MAI-Voice-2 and 32% cheaper, at $15 per 1M characters. > > MAI-Voice-2-Flash is also in public preview and powers Dynamics 365 Contact Center, our enterprise platform for call center agents, and reduces GPU costs up to 89%.… > > — Mustafa Suleyman (@mustafasuleyman) [July 23, 2026](https://x.com/mustafasuleyman/status/2080336147256127960?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) MAI-Voice-2-Flash focuses on speed, scale and lower operating costs. Microsoft says it is twice as fast as MAI-Voice-2 and 32% cheaper while retaining natural prosody and high acoustic quality. It costs $15 per 1 million characters and is intended for responsive voice agents and large call-center workloads. The previews arrive as Microsoft’s broader MAI stack moves deeper into its products. Bing Image Creator now uses MAI-Image-2.5 end-to-end by default, while PowerPoint uses the model for image-to-image work and OneDrive relies on it for key editing tasks. Microsoft reports up to 84% lower GPU costs in PowerPoint compared with GPT-Image-2, plus a 26% rise in save rates and about 25% lower P95 latency in OneDrive. MAI-Voice-2-Flash now powers Microsoft’s enterprise contact-center platform and is integrated into Azure Voice Live, with Microsoft citing GPU cost reductions of up to 89% in the contact-center deployment. The launch follows Microsoft’s year-long push to build purpose-built models trained on clean, traceable enterprise-grade data without distillation from third-party systems. Its strategy is to match each product with a different point on the quality, speed and cost curve while controlling the models that serve millions of users. WPP Global Chief Creative Officer Rob Reilly described the image model’s text rendering as a breakthrough and said its natural-language editing makes creative iteration faster and more intuitive. [Source](https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/?ref=testingcatalog.com) ### ICYMI: Cursor launches Router to cut costs for coding models URL: https://www.testingcatalog.com/icymi-cursor-launches-router-to-cut-costs-for-coding-models/ Last updated: 2026-07-25T22:30:06.000Z Cursor has launched Cursor Router, an intelligent model router for Teams and Enterprise customers that selects an AI model before each coding request runs. Available now on desktop, web, iOS, the Cursor CLI, and Cursor’s SDK, it targets frontier-level performance while reducing spending. Cursor says early-access customers cut costs by roughly 30% to 50%, while online A/B tests across millions of requests showed frontier-quality performance at 60% savings. > Introducing Cursor Router, our intelligent model router that selects the right model for the task at hand. > > Router delivers frontier-quality results at 60% lower cost. [pic.twitter.com/R0YABowFKg](https://t.co/R0YABowFKg?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [July 22, 2026](https://x.com/cursor%5Fai/status/2079993729532989500?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The system uses a classifier trained on more than 600,000 live requests. It examines the query, context, task complexity, domain, and Cursor’s observations of model behavior. Routine work can go to lower-cost models, interface tasks to models chosen for visual taste, and long-horizon problems to frontier reasoning models. Training and production measurements account for cache misses caused by switching models. Users select Auto in the model picker and choose Intelligence, Balance, or Cost. Intelligence aims to match the strongest available models, Balance targets frontier-level quality at a lower price, and Cost prioritizes capability while controlling token spending. Administrators can enable the router by team or group, restrict modes and individual models, and set a default. ![Cursor](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/fable-quality-results-dark-1-.webp) Cursor reports that Auto Intelligence reached satisfaction near Fable at about 60% lower cost and scored roughly 15% above Opus 4.8 at nearly the same cost. Auto Balance exceeded Opus 4.8 at about 36% lower cost and matched GPT-5.6 Sol satisfaction at a lower spending rate. Reported cost per commit was $6.76 for Intelligence and $4.63 for Balance, compared with $12.69 for Fable 5 and $7.34 for Opus 4.8. The company based its findings on online A/B tests measuring user responses and keep rate, or how much agent-generated code remained in a codebase. During two weeks of early access, three high-volume enterprise accounts with thousands of users reportedly saved 30% to 50% without a drop in quality. [Cursor](https://www.testingcatalog.com/tag/cursor/) routes hundreds of millions of coding requests across models and providers each week. With roughly 60% of its developers choosing one daily model, routine tasks can otherwise run at frontier-model prices. Cursor Router addresses that cost gap alongside its tool-calling system, which loads less common tool descriptions only when needed, while Composer handles routine work and Grok 4.5 expands the pool for harder tasks. [Source](https://cursor.com/blog/router?ref=testingcatalog.com) ### Airtap launches text-based AI agent for mobile tasks URL: https://www.testingcatalog.com/airtap-launches-text-based-ai-agent-for-mobile-tasks/ Last updated: 2026-07-25T20:58:48.000Z Airtap has introduced a text-based method to manage everyday phone tasks by launching an iMessage and RCS interface. This innovation transforms an ordinary message thread into a control surface for the apps already installed on a phone. Unlike a chat companion that answers questions, Airtap acts as an agent that operates mobile apps on a user's behalf. It can place orders, clip coupons, check account activity, and post to social platforms, all from within the conversation. The message thread itself serves as the interface, eliminating the need for installation or learning a separate dashboard. 0:00 /0:22 1× Getting started is as simple as sending a single text. Any message sent to the Airtap number automatically creates an account, with no signup form, download, or setup command required beforehand. Airtap reports progress back as plain text as it works. A link appears only when the agent genuinely cannot proceed without input, usually for a first-time sign-in to a third-party service such as a store, streaming, or banking account. Each sign-in link is single-use and scoped to one task. Once the user signs in and returns to the thread, the agent resumes its work independently. SPONSORED Check out the official website! [Test it! ](https://airtap.ai/?ref=testingcatalog.com) The clearest fit for Airtap is with tasks that are entirely within mobile apps and repeat on a schedule. Clipping digital grocery coupons is one example the product focuses on, as supermarket, big box, and pharmacy apps often hide savings behind manually toggled offers that reset weekly. ![Airtap](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Home-Airtap-07-14-2026_05_53_PM.jpg) This task is tedious on a phone and largely absent on desktop. Airtap can work across these apps, clip every available coupon, and repeat the routine weekly without the user opening a single app. This pattern extends to food ordering, class and restaurant booking, and monitoring a limited release until it opens. The agent persists in achieving a goal and escalates until it succeeds, rather than stopping at the first obstacle. The emphasis is on operating real apps that lack usable web equivalents, not on holding a conversation. Airtap is built around a cloud phone model, where the agent navigates a real mobile environment by reading the screen and tapping through app flows as a person would, rather than relying on official APIs that most consumer apps do not expose. This approach allows it to access native apps where a significant portion of daily errands and decisions occur, spanning delivery, fitness, payments, and social feeds. The company has positioned the tool as agent infrastructure for phone-based tasks, with the iMessage and RCS thread serving as the entry point for users who want the capability without technical setup. The product is available via text and automatically recognizes returning users by phone number, ensuring the thread remains continuous across separate tasks. ### Offloop launches AI Agent Workspace for recurrent work URL: https://www.testingcatalog.com/offloop-launches-ai-agent-workspace-for-recurrent-work/ Last updated: 2026-07-23T15:44:45.000Z Offloop has opened its agent workspace in private beta, providing invite-based access to a system that maintains recurring teamwork within persistent Channels instead of restarting it with every prompt. The product is designed for small teams managing launches, customer follow-up, research, feedback, growth, and operations, where progress often stalls not due to hard problems but due to follow-through that no one has time to chase. > Introducing Offloop! > > We're a team of four. Today our multi-agent harness hit state of the art on GDPval, ahead of Claude code and Codex across jobs that pay $2.4 trillion a year in the US. > > Offloop gives every knowledge worker what the Fortune 500 spends billions on: a… [pic.twitter.com/mUVKXscwQA](https://t.co/mUVKXscwQA?ref=testingcatalog.com) > > — Offloop (@Offloop) [July 23, 2026](https://x.com/Offloop/status/2080309689670361172?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The unit of work is the Channel. People, agents, messages, files, schedules, signals, tool progress, and finished artifacts are all located in the same place, allowing an agent picking up a task to inherit the context, owners, and permissions that preceded it. Schedules and incoming signals can independently activate an agent, transforming a one-off request into a loop that continues running. In one setup demonstrated by the company, an agent is granted access to a Google Workspace inbox and subsequently activates the Channel with a summary whenever a matching email arrives. Actions remain visible throughout, and individuals can inspect, redirect, approve, or stop a run at any point. 0:00 /0:23 1× Offloop Coordinating several agents within one Channel is managed by D1, a dispatcher model developed by Offloop for this purpose. With multiple agents sharing a workspace, a decision-maker is needed to determine which agent acts at each step and to respond when context shifts or new updates occur, and D1 fulfills this role. Teams can connect their own AI subscription account or utilize inference provided by Offloop. Connectors currently include Notion, GitHub, GitLab, Sentry, Supabase, Vercel, Railway, Cursor, Codex, and ChatGPT, with a macOS desktop app accompanying the workspace. 💡 Use code 6M50DJ and get access to the closed beta! Offloop is also quantifying the multi-agent framework. According to figures released by the company, its multi-agent system scores 84.9 percent on GDPval and 67.2 percent on JobBench, surpassing the frontier multi-agent baselines it tested against, at approximately a fifth of the cost per task. GDPval is OpenAI's benchmark of real-world economically valuable work spanning 44 occupations across nine sectors, while JobBench, published by researchers at the University of Washington, evaluates agents on the tasks professionals wish to delegate. Offloop presents the result as the first multi-agent system to demonstrate its value on real-world benchmarks. SPONSORED Start testing Offloop [Learn more ](https://offloop.org/invite?code=6M50DJ&ref=testingcatalog.com) The company behind the product is Intelligence Software, Inc., and Offloop states that the team previously developed general-purpose agents at scale and recognized where a single agent reaches its limits. Access is available across three tiers: an invite-only Operator plan for individuals converting recurring work into agent-assisted Channels, a Team Pilot plan that adds shared Channels, team memory, connector setup, and priority onboarding, and an enterprise tier that includes custom connectors, SSO planning, and an audit-friendly runtime. Invite codes for the beta are being distributed through the Offloop site. ### Anthropic preparing for potential Claude Opus 5 rollout URL: https://www.testingcatalog.com/anthropic-preparing-for-potential-claude-opus-5-rollout/ Last updated: 2026-07-23T10:46:59.000Z Anthropic looks close to shipping Claude Opus 5, and the signals now point **past internal testing** into limited partner hands. The sharpest artifact so far is a research entry that surfaced briefly in a coding tool's model picker earlier this month before being pulled within hours. It was listed as an early access preview with per-turn controls, safety fallbacks, a one million token context window, and an extra-high-effort setting. A matching model string was then reported on Google's Vertex quotas catalog, echoing the sequence that preceded Sonnet 5\. No pricing, system card, or public model identifier exists yet, so the spec list stays provisional. > Early preparations for Claude Opus 5 have been spotted again. Seems imminent. > > — M1 (@M1Astra) [July 23, 2026](https://x.com/M1Astra/status/2080092848926392525?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Should it land, the model would appear in the [Claude](https://www.testingcatalog.com/tag/claude/) apps model selector for paid tiers, in Claude Code, on the Claude Platform, and across the three cloud partners, most likely replacing Opus 4.8 rather than running beside it. The fallback detail is the telling one: routing flagged prompts down to 4.8 is the same arrangement used above the Opus tier, which hints that this release is being positioned nearer the frontier than its number suggests. That positioning matters because Opus has been squeezed from both sides. Sonnet 5 arrived at the end of June performing close to 4.8 at a fraction of the cost, while Fable and Mythos sit above it. Subscription-inclusive Fable 5 access ended on 19 July with usage credits taking over. A stronger Opus would give Max, Team, and Enterprise customers somewhere to move without paying Mythos rates, which is the practical reason timing chatter has settled on this week. [Join Dev Mode ](https://discord.com/invite/devmode?ref=testingcatalog.com) Cadence supports the guess without proving it: 1. 70 days from 4.6 to 4.7 2. 42 from 4.7 to 4.8 3. 56 so far since Expectations run toward a clear step up on coding work, though Sonnet 5's reception is a reminder that they are not always met. Anthropic has said nothing. ### OpenAI tests Codex Realtime Voice Mode for ChatGPT URL: https://www.testingcatalog.com/openai-tests-codex-realtime-voice-mode-for-chatgpt/ Last updated: 2026-07-23T10:07:20.000Z OpenAI appears to be building a [real-time voice mode for Codex](https://www.testingcatalog.com/openai-prepares-real-time-voice-mode-for-pets-in-codex/) that extends well beyond code. Recent additions to the ChatGPT bundle include strings and supporting functionality that describe a general-purpose agentic assistant you can communicate with via phone. The voice layer maintains the conversation while worker agents perform tasks in the background. **The system prompts behind it reference checking Slack, pulling up documents and calendars, controlling Spotify, browsing, shopping, and food ordering, including scanning Uber Eats for dinner.** This represents a personal assistant framework rather than a developer one, and it operates on top of Codex Remote Control. Thus, voice would become a means to dispatch and manage several tasks simultaneously on a laptop, with results read back to the user. > [pic.twitter.com/1kEPJoo0Sg](https://t.co/1kEPJoo0Sg?ref=testingcatalog.com) > > — M1 (@M1Astra) [July 23, 2026](https://x.com/M1Astra/status/2080163439511588948?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Several aspects remain uncertain. It is unclear which model will handle the conversations, with the bidirectional voice model spotted earlier this summer being the obvious candidate, though nothing has been confirmed yet. It is also uncertain whether Codex Real-Time will have its own entry in the voice model selector. A real-time voice section was observed inside the Codex app before the work was restructured more broadly, and it has not appeared on the ChatGPT desktop app, possibly residing in a bundle that has not been released. > Unbelievable excited for what’s coming together. Tomorrow is feeling codexy > > — Tibo (@thsottiaux) [July 23, 2026](https://x.com/thsottiaux/status/2080144499716800513?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The framework aligns with the company's direction. OpenAI recently [integrated the Codex app into the ChatGPT desktop app](https://www.testingcatalog.com/openai-launches-chatgpt-work-for-pro-enterprise-and-edu-plans/), positioning Codex alongside Chat and Work. Codex Remote reached general availability across paid plans in June. A spoken layer that covers both coding threads and everyday tasks suits a company that is consolidating its interfaces into one assistant, potentially attracting users who never open a terminal. [Join Dev Mode ](https://discord.com/invite/devmode?ref=testingcatalog.com) The timing is uncertain, and this may be early groundwork. Codex staff have hinted at something for tomorrow, with discussions pointing to parallel worker agents and faster Cerebras-served inference for GPT-5.6 Sol, rather than a new silicon deal, since that partnership has been ongoing since January and has already produced Codex-Spark. Whether the voice feature will be included in the same release is the aspect worth monitoring. ### SpaceXAI develops deployable applications for Grok Build URL: https://www.testingcatalog.com/spacexai-develops-deployable-applications-for-grok-build/ Last updated: 2026-07-23T00:56:41.000Z xAI is preparing to push Grok's coding surface beyond generation and into distribution. Work is underway on functionality that would let users deploy, host, and publicly share apps built with Grok and Grok Build, an approach that closely mirrors what Google AI Studio already offers through its build-and-publish flow. xAI also appears to be developing support for custom domains, letting a shared app live on an address of the creator's choosing rather than a generic link. Neither capability is available yet, and no timeline has surfaced, but the direction lines up with earlier community reports pointing to deployment support across dozens of auto-detected frameworks. If shipped, the pipeline would mostly serve hobbyist builders and small teams who want to go from prompt to a working URL without touching hosting infrastructure, and it would recast the Grok app as a publishing platform rather than a code generator alone. > Check out GrokCraft on [https://t.co/YK6T9wAJld](https://t.co/YK6T9wAJld?ref=testingcatalog.com) built using the [https://t.co/9ZJ2Pejnm1](https://t.co/9ZJ2Pejnm1?ref=testingcatalog.com) App Builder! > > good find [@Oldmannotme](https://x.com/Oldmannotme?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com)! [https://t.co/eo2Bv0SvMs](https://t.co/eo2Bv0SvMs?ref=testingcatalog.com) [pic.twitter.com/iAZbkVDYcn](https://t.co/iAZbkVDYcn?ref=testingcatalog.com) > > — ️ (@blankspeaker) [July 22, 2026](https://x.com/blankspeaker/status/2080006022073344433?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Additionally, since June, paid [Grok](https://www.testingcatalog.com/tag/grok/) plans draw from a single weekly usage pool shared across Chat, Imagine, Voice, and Build, with extra credit top-ups sold on the web once the pool runs dry. That top-up flow now looks set to reach the Grok mobile app as well, removing a detour through the website for anyone who hits their limit mid-session on a phone. > SPACEXAI 🔥: The next 2T params Grok model is expected to finish training next week and supposed to exceed Kimi K3 in performance with a better speed and token efficiency. > > If we will be able to see this model in September, that would mean that SpaceXAI managed to put this… [https://t.co/4Y7Lc4JVkV](https://t.co/4Y7Lc4JVkV?ref=testingcatalog.com) [pic.twitter.com/bP1xztCOWI](https://t.co/bP1xztCOWI?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 18, 2026](https://x.com/testingcatalog/status/2078444145819984100?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model roadmap anchors all of this. Elon Musk said on July 18 that xAI's 2 trillion parameter base model, reportedly named Grok 4.6, will finish initial training next week. He claims it beats the 1.5 trillion parameter foundation behind Grok 4.5 in every way and might exceed Moonshot's Kimi K3 while keeping speed and token efficiency close to Grok 4.5\. With customer availability targeted for August, a stronger coding model could arrive just as the distribution layer around it takes shape. ### Claude Voice Mode to get Opus and Sonnet model options URL: https://www.testingcatalog.com/claude-voice-mode-to-get-opus-and-sonnet-model-options/ Last updated: 2026-07-23T00:07:22.000Z Anthropic looks set to give Claude's voice mode a real upgrade, and the clearest signal yet landed this week. A [model selector has sat inside the voice interface](https://www.testingcatalog.com/claude-code-managed-agents-and-model-selector-for-voice-mode/) for roughly three weeks, but until now the choice was cosmetic: whichever option you picked, the session still ran on Claude Haiku 4.5, a model that has not been refreshed in some time. That shifted today. **Selecting Opus or Sonnet now routes voice conversations through those models** rather than quietly falling back to Haiku, which points to an imminent public rollout, probably within days (\*This feature is not yet available to the public and is currently hidden under a feature flag). 0:00 /1:34 1× Opus and Sonnet on Claude Voice Mode The picker currently exposes Opus, Sonnet, and Haiku, with a control for setting the thinking level beside it. Claude Fable, the company's Mythos-class model, is absent from voice for now. The pipeline stays text-to-speech rather than a speech-native model: Claude handles the reasoning and drives the conversation, while spoken output appears to lean on ElevenLabs in the background, the same subcontractor listed since the feature launched. Anyone who found Haiku too thin for real spoken work stands to gain the most. Routing Opus or Sonnet into voice opens tasks that were awkward before, from longer reasoning to tool-heavy requests handled by speech. In testing, interruption handling held up: the model pauses the moment it hears you and resumes sensibly, and it tolerated long silences without cutting in early, so the mode stays usable despite the added intelligence. The direction says something about [Anthropic's](https://www.testingcatalog.com/tag/claude/) strategy. Where OpenAI is pouring resources into a [bidirectional](https://www.testingcatalog.com/openai-rolls-out-gpt-live-voice-for-chatgpt-on-web-and-mobile/), speech-native model, Anthropic is taking the opposite path, placing its strongest reasoning models on a conventional voice stack rather than chasing voice-specific tricks. No formal post has appeared, though the functional switch behind the selector suggests one is close. ### Anthropic develops Claude-driven Managed Projects URL: https://www.testingcatalog.com/anthropic-develops-claude-driven-managed-projects/ Last updated: 2026-07-22T10:31:52.000Z Anthropic appears to be preparing a new kind of project for Claude, one where the assistant does more than hold files and answer questions. A recent build surfaces a choice, at the moment a project is created, between a standard project and a "managed" one, described with the line that **Claude takes on tasks and keeps the project organized**. That framing points to a persistent Claude environment that carries the surrounding context, works with some autonomy, and could run scheduled work rather than waiting for each prompt. Beside it, the [projects area of the side navigation](https://www.testingcatalog.com/anthropic-tests-new-placement-for-projects-on-claude-desktop/) is being reworked to list only pinned projects, based on tester feedback. Both are marked internal for now, limited to staff and a small circle of trusted testers, with no release date attached. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Screenshot-2026-07-22-at-01.16.48.png) Internal Project section update ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Screenshot-2026-07-22-at-01.17.29.png) Managed Project creation screen The pieces line up with work Anthropic already ships elsewhere. Its developer platform has offered Claude Managed Agents since April, cloud-hosted agents that hold state across sessions and refine their own memory through a scheduled process the company calls dreaming. Cowork projects already run scheduled tasks, and Claude Code carries context between sessions through project memory and instructions. Managed projects read like those threads gathered into one consumer surface, where sessions in a project would share memory and instructions and each one feeds the project's wider knowledge. The build also notes a project can be personal or shared, with shared projects held for Team and Enterprise plans. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Claude-Code-07-22-2026_02_13_AM.jpg) Projects on Claude Code For Anthropic, this fits a steady move from chat toward standing, self-organizing workspaces. The company has tried knowledge bases before and has been running Conway, its always-on internal agent, which is **set to close on July 24**, timing that could mark a handoff rather than an ending. If managed projects reach users, they would offer a project that quietly maintains itself between visits, a shape no rival lab currently sells. A firm date is still missing, but internal testing is plainly live. ### Google launches Gemini 3.6 Flash and Gemini 3.5 Flash Lite URL: https://www.testingcatalog.com/google-launches-gemini-3-6-flash-and-gemini-3-5-flash-lite/ Last updated: 2026-07-22T10:31:26.000Z Google has released new additions to its Gemini Flash model series, targeting developers and enterprise customers building AI agents that require high efficiency, reduced latency, and reliability in production. The company is launching Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, each tailored for specific use cases such as coding, multimodal tasks, and cybersecurity, respectively. 3.6 Flash and 3.5 Flash-Lite are available immediately through the Gemini API and Gemini Enterprise, while 3.5 Flash Cyber will be accessible to governments and trusted partners via CodeMender in an upcoming limited pilot. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/gemini-3-6-flash__evals__figure-.width-2000.format-webp.webp) Gemini 3.6 Flash demonstrates a 17% reduction in token usage compared to its predecessor, 3.5 Flash, and achieves lower costs per output token. It shows improved precision in coding, knowledge work, and multimodal tasks, including document parsing and data analysis. Safety safeguards are strengthened, especially in domains like CBRN and cyber offense, reducing risks of misuse. Gemini 3.5 Flash-Lite offers the fastest output in the series at 350 tokens per second and is priced for high-volume, cost-conscious workflows. It outperforms earlier Flash-Lite models in coding and agentic tasks. Gemini 3.5 Flash Cyber is fine-tuned for vulnerability detection and patching, and its deployment is intentionally restricted to prevent misuse. > As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models. > > Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️ > > Built on top of 3.5 Flash, in… [pic.twitter.com/8N63Ws6laR](https://t.co/8N63Ws6laR?ref=testingcatalog.com) > > — Google AI (@GoogleAI) [July 21, 2026](https://x.com/GoogleAI/status/2079617029473182132?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > 3.5 pro is testing with partners! will hopefully land soon. > > — Logan Kilpatrick (@OfficialLoganK) [July 21, 2026](https://x.com/OfficialLoganK/status/2079596415509303596?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) In case you are waiting for Gemini 3.5 Pro as well 👀 Google continues to iterate on the Gemini model family, with 3.5 Pro in partner testing and Gemini 4 in pre-training, reflecting its ongoing focus on AI agent infrastructure and safety. [Source](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/?utm%5Fsource=x&utm%5Fmedium=social&utm%5Fcampaign=&utm%5Fcontent=blog.google/innovation-and-ai/%E2%80%A6/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber?%E2%80%A6%3E) ### Perplexity tests OpenRouter integration for Computer URL: https://www.testingcatalog.com/perplexity-tests-openrouter-integration-for-computer/ Last updated: 2026-07-21T14:38:34.000Z Perplexity appears to be building a bridge between Perplexity Computer and OpenRouter, one that would let subscribers connect their own OpenRouter account and route the agent through the wide catalog of models hosted there. The wiring sits inside the product but does not yet function, pointing to an early internal test rather than anything close to a public switch. The appeal is mostly about cost. Perplexity Computer coordinates a large roster of models to plan and carry out multi-step work across local and remote environments, and that orchestration burns through credits quickly, the practical ceiling that keeps the agent out of reach for many. Routing through OpenRouter would open far cheaper options, and in some cases free ones, loosening the constraint that has defined who can realistically lean on Computer day to day. Power users have already been asking for this kind of account-level model connection. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Perplexity-07-21-2026_01_14_AM--1--1.jpg) Where it would live is easy to picture: a connections panel to link the key, then a model layer feeding [Computer's](https://www.testingcatalog.com/perplexity-released-personal-computer-to-all-max-subscribers/) routing rather than the credit pool [Perplexity](https://www.testingcatalog.com/tag/perplexity/) meters today. Pro subscribers gained Computer earlier this year, with Max holding higher limits and a monthly credit grant, so a route around that metering would reshape the value calculation across both tiers. > Breaking w/ [@KevKubernetes](https://x.com/KevKubernetes?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [@validapau](https://x.com/validapau?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com): > > OpenRouter, which helps devs choose between different models, has discussed a potential sale to a bigger tech company that could value it at billions of dollars, a steep premium to its $1.3B valuation. > > More here:[https://t.co/FdDj0UoRDv](https://t.co/FdDj0UoRDv?ref=testingcatalog.com) > > — Stephanie Palazzolo (@steph\_palazzolo) [July 17, 2026](https://x.com/steph%5Fpalazzolo/status/2078202590349676919?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Perplexity & OpenRouter? Is there a chance for these two to get together? For a company built on multi-model choice, a single research surface fronting Claude, GPT, Gemini, and others, handing users a personal pipe to an even larger pool fits the strategy cleanly. It also answers a run of complaints about tightened usage caps and upgrade prompts. If it ships, the move would widen Computer's audience well beyond the paying core it serves now, though nothing is confirmed and no timeline has surfaced. ### Anthropic set to end Conway test as wider preview expected soon URL: https://www.testingcatalog.com/anthropic-set-to-end-conway-test-as-wider-preview-expected-soon/ Last updated: 2026-07-20T15:14:41.000Z Anthropic's [Conway](https://www.testingcatalog.com/exclusive-anthropic-tests-its-own-always-on-conway-agent/) experiment is approaching a fork in the road. A notice we recently spotted on the Conway interface states that the project will be discontinued on July 24th and asks testers to export their data before then. Conway, the always-on agent Anthropic has been running internally with employees and, by some indications, a small friends-and-family circle, first surfaced this spring as a standalone instance able to run Claude Code, operate a browser, send notifications, and wake up via webhooks. > Conway access ends Fri July 24 at 5 pm PT. To export your data, ask Conway: "export my data". The notice reads two ways. The first is a full wind-down: **Anthropic may have concluded that a remote container where Claude works in an isolated environment does not fit its roadmap**, and that its agent bets are better placed on the desktop path already shipping through Cowork, where Claude acts as a personal agent on the user's own machine. In that case, Conway would likely vanish from the codebase without ever being announced, closing one of the company's most intriguing side projects. The second reading is that **internal testing is wrapping up because a broader phase is next**. Anthropic tends to open new capabilities to Max subscribers first, and a limited preview along those lines would fit the pattern. Under this scenario, Conway would become reachable across Anthropic's apps as a remote container where users install plugins and skills, hand off different types of work, and configure webhooks that fire when outside events occur. Alongside it, a standard referred to as UI tabs is expected to let users define and share their own component panels, a custom Conway control surface, for instance, building on the extension format seen in earlier findings. Which way this breaks matters beyond one experiment: it would show whether [Anthropic](https://www.testingcatalog.com/tag/claude/) believes persistent agents belong on the user's device or in the cloud, a question rival labs are circling with always-on projects of their own. By the end of July, the answer should be on the record. ### Anthropic tests Penlight for live clinical transcripts and AI research URL: https://www.testingcatalog.com/anthropic-tests-penlight-for-live-clinical-transcripts-and-ai-research/ Last updated: 2026-07-20T12:57:33.000Z Anthropic is developing a new Claude product called **Penlight**, designed to accompany clinicians during patient visits. It records conversations and transforms them into live, searchable AI sessions. According to code and an early interface reviewed by TestingCatalog, Penlight is intended to enhance clinical interactions by providing real-time transcription and interaction capabilities. Penlight appears as a separate item under the Products section of Claude’s sidebar, alongside tools such as Claude Code, Design, and Conway. It features its own home screen and dedicated session pages, making it a standalone Claude experience rather than another mode inside the regular chat interface. The product is currently hidden behind an internal feature flag, and its early interface is unfinished. Several labels are delivered through gated localization messages that were unavailable during testing, resulting in blank buttons and badges on the home screen. The Penlight name is confirmed by the sidebar and page-heading code. Once a user begins a session, Penlight requests microphone access and starts streaming audio to a dedicated Anthropic service. The browser encodes the recording as mono Opus audio at 16 kHz, a configuration optimized for speech rather than high-fidelity playback. The service produces a live transcript that can identify different speakers. Users can show or hide the transcript while continuing to interact with Claude in a parallel chat panel, allowing them to ask questions without leaving the ongoing visit. Penlight also tracks whether a recording is active on another device and offers a way to resume interrupted live sessions. The underlying interface goes beyond basic transcription. Research results are structured around scientific articles and can include the journal, authors, publication year, DOI, and PubMed links. This suggests Penlight is intended to help users check clinical questions against medical literature while a conversation is still underway. Every session is internally categorized as a “visit.” The service model also includes a field for a generated note, although the current interface does not appear to display that note yet. Together, these details point toward an ambient clinical assistant that could eventually transcribe an appointment, answer questions using medical sources, and prepare documentation for review. The product’s home screen is built around a session list, with sections for recent activity and sessions that “need” the user’s attention. A persistent control can return the user to an active recording. This indicates Anthropic is designing Penlight as an ongoing workflow for managing multiple visits, not merely a one-off voice conversation with Claude. Penlight aligns closely with Anthropic’s broader [healthcare](https://www.anthropic.com/news/healthcare-life-sciences?ref=testingcatalog.com) push. In January, the company expanded Claude for Healthcare and explicitly identified ambient scribing, chart review, and clinical decision-support tools as potential applications. It also added access to sources such as PubMed through its healthcare offering. Anthropic’s healthcare materials describe Claude as infrastructure for healthcare companies building those experiences. Anthropic’s [healthcare product page](https://claude.com/solutions/healthcare?ref=testingcatalog.com) already demonstrates a workflow in which Claude turns a patient-visit recording into a clinical summary, assessment, and plan. Penlight appears to bring a version of that workflow directly into the main Claude application, with recording and assistance happening live rather than after an audio file is uploaded. Important questions remain. The code does not establish whether Penlight is intended for individual clinicians, healthcare organizations, or an internal research program. There is also no visible information about EHR integrations, regional availability, compliance requirements, or pricing. Anthropic could substantially change or abandon the product before release. Still, Penlight is one of the clearest signs yet that Anthropic is exploring first-party, industry-specific applications inside Claude. Instead of only supplying models to ambient-scribing companies, the company appears to be testing a clinical workflow of its own, one that listens during a visit, keeps Claude available in the room, and connects the conversation directly to medical research. ### Anthropic tests new placement for Projects on Claude Desktop URL: https://www.testingcatalog.com/anthropic-tests-new-placement-for-projects-on-claude-desktop/ Last updated: 2026-07-20T08:40:52.000Z Anthropic appears to be preparing another round of layout changes for the [Claude](https://www.testingcatalog.com/tag/claude/) desktop app, barely two weeks after Chat and Cowork moved into a shared Home. Work surfacing in recent builds shows the toggle between the two modes leaving the prompt bar, where it landed with the [July 7 update](https://www.testingcatalog.com/anthropic-brings-claude-cowork-to-web-and-mobile-for-max-users/), and moving to the top of the window. The placement mirrors the merged ChatGPT desktop app, where OpenAI put its mode switcher in the top-left corner after folding the Codex app into [ChatGPT on July 9](https://www.testingcatalog.com/openai-launches-chatgpt-work-for-pro-enterprise-and-edu-plans/). Given how fresh the current design is, this looks like a candidate for an A/B test with a small group of users rather than an imminent rollout. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Claude-Code-07-19-2026_12_30_AM.jpg) The second change concerns projects. Anthropic is testing a way to surface them directly in the navigation sidebar, above recent conversations, with options to pin them, sort them, and decide what shows up there. OpenAI shipped a similar treatment for ChatGPT in June, letting users pin chats and projects and regroup the Recents list, and demand for that kind of sidebar control had been building for a long time. For people juggling many workspaces in Claude, having projects one click away instead of behind a separate page would be a small but meaningful shift. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Claude-Code-07-19-2026_12_30_AM--1-.jpg) Claude Code is getting attention as well. On the web, a dedicated projects section is in the works, where users would be able to create new projects and customize them. Today, Claude Code scopes its work to a single folder or repository, so a management layer on top could give desktop and web users a proper overview of everything Claude is building for them. Taken together, the changes fit Anthropic's current push to make the desktop app the center of gravity for deep work while Cowork expands to web and mobile in beta. Whether any of this ships is another question, though, as Anthropic is known for experimenting with interface ideas and quietly rolling them back when they underperform. ### Google plans to upgrade Gemini Enterprise connectors URL: https://www.testingcatalog.com/google-plans-to-upgrade-gemini-enterprise-connectors/ Last updated: 2026-07-19T00:05:37.000Z Google appears to be preparing another round of changes for Gemini Enterprise, its front door for AI-assisted work across company data. In recent builds, the platform's main prompt bar has picked up the glowing treatment already familiar from the consumer Gemini app, bringing the two products close to visual parity. It is a small touch, but one that fits Google's broader push to make its enterprise assistant feel like the Gemini that hundreds of millions of people already use. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Gemini-Enterprise-07-15-2026_05_34_PM.jpg) More consequential is a new notice telling users that their Google Workspace connectors are being upgraded and may need to be reauthorized to keep working. The message does not spell out what changed under the hood, but the timing is telling. The connector catalog in Gemini Enterprise has grown quickly in recent months, covering Slack, Microsoft Teams, Notion, Linear, Jira, ServiceNow, and dozens of other workplace tools, with new actions such as creating pages and updating tickets landing steadily. A reauthorization wave usually points to expanded permissions or reworked plumbing, and the likely payoff is more dependable behavior whenever Gemini reads from or writes to those services. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Gemini-Enterprise-07-15-2026_05_35_PM.jpg) [Gemini Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/), the always-on agent [Google](https://www.testingcatalog.com/tag/gemini/) introduced at I/O in May, now appears behind a feature flag inside Gemini Enterprise, though it remains hidden for the moment. Google has said Spark would reach Gemini Enterprise customers soon and that the agent will be able to work through existing connectors such as SharePoint, OneDrive, and ServiceNow, so a background worker acting across a company's tools is only as capable as the connections beneath it. Seen that way, the connector overhaul looks like groundwork rather than housekeeping. Whether any of this has reached all tenants is unconfirmed, and Google typically ships such changes gradually, so some organizations may already see pieces of the update while others wait. ### Google preparing Gemini Live and Skills for web rollout URL: https://www.testingcatalog.com/google-preparing-skills-and-gemini-live-for-web-rollout/ Last updated: 2026-07-18T14:21:12.000Z Google appears to be preparing two notable additions to Gemini on desktop and web. In recent builds, a Gemini Live button has surfaced in the web version, suggesting the real-time voice mode may finally move beyond mobile. Live remains absent even from the desktop app today, though testing is reportedly underway within the scope of Google's [Trusted Tester](https://www.testingcatalog.com/google-tests-voice-dictation-and-magic-pointer-on-gemini-desktop/) program, and a simultaneous debut across desktop and web looks plausible. Google itself promised new voice capabilities for the macOS app at I/O in May, and earlier signs in that app pointed to voice selection options and a screen-sharing overlay, so the groundwork has been visible for months. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Screenshot-2026-07-16-at-23.54.21.webp) The same builds also carry traces of skills being wired into regular chats. Skills currently live exclusively inside Gemini Spark, the autonomous agent Google gates behind its AI Ultra subscription, where they act as reusable instruction packages the agent applies to recurring tasks. The traces suggest ordinary chat users could soon: 1. Upload their own skills 2. Pick from predefined ones 3. Build new ones from scratch 4. Have Gemini generate a skill on their behalf, even without Spark access > Google is working on a native menu for Skills on Gemini desktop. These Skills may become available in all chats (fingers crossed). > > Users will be able to upload, create, and edit their Gemini Skills, as well as use Gemini to create them. > > Skill folders will also be available… [pic.twitter.com/ZEmNAdt5wE](https://t.co/ZEmNAdt5wE?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 16, 2026](https://x.com/testingcatalog/status/2077879966889382371?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) That would mirror what Claude and ChatGPT already offer, closing one of the more visible feature gaps in Google's assistant, and it fits a broader pattern: Chrome received a lighter prompt-shortcut take on skills in April, and Gemini Enterprise already lets workers invoke skills mid-chat. Timing is the open question. [Gemini 3.6 Flash](https://www.testingcatalog.com/google-might-be-testing-gemini-flash-upgrade-on-lm-arena/) appears to be in preparation while Gemini 3.5 Pro has slipped repeatedly, most recently past a mid-July target, so a stopgap model could arrive first. Something may surface on Google by the end of July, but nothing here has been announced, and plans of this kind can shift before launch. Together with Live on web, chat-level skills would make the everyday Gemini experience considerably more capable. ### OpenAI unveils GPT-Red to boost AI safety with internal red-teaming URL: https://www.testingcatalog.com/openai-unveils-gpt-red-to-boost-ai-safety-with-internal-red-teaming/ Last updated: 2026-07-17T21:39:50.000Z OpenAI has unveiled GPT‑Red, its strongest automated safety red-teaming model, built to discover vulnerabilities in AI systems before deployment. The internal-only model targets prompt injections, including malicious instructions hidden inside emails, webpages, local files, code repositories, and tool outputs. GPT‑Red learns through self-play reinforcement learning. It attacks a group of defender models across realistic scenarios while both sides are trained simultaneously. GPT‑Red receives rewards for causing valid failures, while defenders are rewarded for resisting attacks and completing the original task. As defenders grow more robust, GPT‑Red must develop stronger and more diverse attack methods. > OpenAI announced GPT-Red, an internal model for finding prompt-injection vulnerabilities at scale. > > \> GPT‑Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt injection attacks. > > \> We use GPT‑Red to adversarially train GPT‑5.6, making it much… [https://t.co/GYbq0ob8TL](https://t.co/GYbq0ob8TL?ref=testingcatalog.com) [pic.twitter.com/ufmqqmM3cL](https://t.co/ufmqqmM3cL?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 15, 2026](https://x.com/testingcatalog/status/2077513154112929909?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) OpenAI trained the system using compute comparable to some of its largest post-training runs. In tests involving previously unseen scenarios, GPT‑Red successfully attacked GPT‑5.1 in 84% of cases, compared with 13% for human red-teamers. It could also break nearly every tested OpenAI model through GPT‑5.5. The model demonstrated attacks against real agentic systems. It manipulated an AI-operated vending machine to reduce product prices to $0.50, order an item worth more than $100 at that price, and cancel another customer’s order. In separate tests, it caused a Codex CLI agent powered by GPT‑5.4 mini to exfiltrate sensitive data across a custom set of scenarios. OpenAI has incorporated attacks generated by GPT‑Red into the training of production models since GPT‑5.3\. GPT‑5.6 Sol recorded six times fewer failures on the company’s hardest direct prompt-injection benchmark than its leading production model from four months earlier. Its failure rate against GPT‑Red’s direct prompt injections fell to 0.05%. An earlier GPT‑Red version also discovered “Fake Chain-of-Thought” attacks, which succeeded more than 95% of the time against GPT‑5.1\. That rate dropped below 10% with GPT‑5.6 Sol. OpenAI reports that these robustness gains did not reduce general model capabilities or cause broader refusal behavior. GPT‑Red will remain separate from public production models because it was deliberately trained with malicious capabilities. It is not being released to ChatGPT users or API developers. OpenAI plans to continue scaling the system alongside human and third-party red-teaming, layered safeguards, and real-time monitoring, using attacks found by current models to train more robust future GPT releases. [Source](https://openai.com/index/unlocking-self-improvement-gpt-red/?ref=testingcatalog.com) ### Perplexity launches SPACE runtime for AI agent tasks URL: https://www.testingcatalog.com/perplexity-launches-space-runtime-for-ai-agent-tasks/ Last updated: 2026-07-17T21:38:25.000Z Perplexity has introduced SPACE, short for Sandboxed Platform for Agentic Code Execution, a new runtime designed to support long-running AI agents that execute code, edit files, and complete multi-step tasks over hours, days, or even months. The company began deploying SPACE as the sandbox layer for Perplexity Computer in June. It now powers 100% of Computer sessions and has already handled millions of sandbox creations and tens of millions of reconnections. Perplexity plans to expand the system across more products and environments, including Linux microVMs, Windows guests, and users’ local machines. > SPACE separates the session from the sandbox running it. > > Each task gets a disposable Firecracker microVM that is destroyed when the work ends. Rolling snapshots preserve live memory and files, so the session can pause, resume, or branch across sandboxes. [pic.twitter.com/uDZ3OUxgXC](https://t.co/uDZ3OUxgXC?ref=testingcatalog.com) > > — Perplexity (@perplexity\_ai) [July 15, 2026](https://x.com/perplexity%5Fai/status/2077432552651141532?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Each SPACE sandbox runs inside a virtual machine with its own guest kernel, separating untrusted workloads from the host and other tenants. A dedicated in-guest daemon controls filesystem and process operations through a private channel, while a network gateway governs outbound traffic. Credentials remain outside the sandbox and can be injected at the network layer or entered by a browser agent. Enterprise customers can also use their own encryption keys for externally stored data. SPACE supports persistent sessions, pausing, resuming, suspension, restoration, rollback, forking, and crash recovery. Regular disk snapshots preserve filesystem changes, while full snapshots capture the paused virtual machine. Suspended sessions are uploaded to object storage, allowing another node to restore the agent with its prior filesystem and running state. Perplexity built the storage layer on btrfs, using copy-on-write cloning and atomic snapshots to reduce startup time and storage consumption. Warm pools keep commonly used templates ready for incoming sessions. In production testing, median sandbox creation time dropped from 185 milliseconds to 60 milliseconds, while 90th-percentile latency fell from 447 milliseconds to 89 milliseconds. This makes SPACE between 3.1 and 5 times faster than Perplexity’s previous sandbox system. [Perplexity](https://www.testingcatalog.com/tag/perplexity/) is positioning SPACE as shared infrastructure for its agent products and future developer-facing runtimes. The platform addresses a central challenge for autonomous agents: preserving state and access to working files while isolating potentially compromised code, credentials, and network activity. [Source](https://research.perplexity.ai/articles/making-space-secure-and-efficient-runtimes-for-long-running-agents?ref=testingcatalog.com) ### NoimosAI unveils Advisor for automated marketing tasks URL: https://www.testingcatalog.com/noimosai-unveils-advisor-for-automated-marketing-tasks/ Last updated: 2026-07-17T16:30:54.000Z AGOS LABS is rolling out Advisor, a new layer inside NoimosAI that moves the platform from executing work on command to proposing the work itself. The feature goes live on July 17 and is built around what the company calls an "approve to execute" workflow. The agent reads the marketing data already connected to a workspace, measures the distance between current performance and the goals set for the account, and surfaces the specific next actions that close that gap. Approval takes a single click, and the task runs. > Every AI assistant waits for you to ask. > > NoimosAI Advisor doesn't. > > It analyzes performance and market insights, finds what works or doesn't, and drafts the work. > > You approve. The work is finished. [pic.twitter.com/Qf1sUu5aVm](https://t.co/Qf1sUu5aVm?ref=testingcatalog.com) > > — NoimosAI (@noimos\_ai) [July 17, 2026](https://x.com/noimos%5Fai/status/2078153138461405518?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The data the Advisor reads is the account's own rather than a generic market signal. Recent social posts, Google Search Console, and Google Analytics feed into the picture the agent builds of where a brand actually stands, which ties each proposed action to observed performance rather than to a template. NoimosAI describes the loop as real-time, with the platform pulling from connected apps and premium data sources on a daily refresh cycle. The framing aimed at the target user is blunt: the actions worth taking are queued and waiting in the morning, and clicking one moves the work forward without a prompt being written. Advisor lands on a platform that already had one-click approval, though positioned later in the process. Until now, the Feed collected finished agent output for review after a human had assigned the task through Chat or set a trigger. Advisor inverts that order by generating the assignment itself. It sits alongside an agent roster covering growth metrics and strategy, competitor strategy, social listening, industry news, social, SEO, GEO, event outreach, media outreach, and conversion rate optimization, and draws on the same memory and knowledge base layer that carries brand context between tasks. SPONSORED Start testing with NoimosAI [Learn more ](https://noimosai.com/?ref=testingcatalog.com) NoimosAI is built by AGOS LABS TECHNOLOGIES LTD, a Dubai-registered company founded by Kosuke Yokoyama, who ran the product through a closed beta before opening it to the public on June 10\. The platform is billed as the world's first all-in-one autonomous AI marketing team, covering strategy planning, execution, and continuous improvement, and is aimed at founders, freelancers, creators, marketers, and small-business operators who want expert-level output without the headcount. Pricing starts at 99 dollars per user per month on the Pro tier, rising to 249 dollars on the Team tier and 499 dollars on the Advanced tier, with a free trial across all plans. Advisor answers the gap that the company's own positioning leaves, since an autonomous marketing team is only autonomous once it decides what to work on, and the approval click is the part the operator keeps. ### Early look at Kimi K3 generations from Moonshot AI on Arena URL: https://www.testingcatalog.com/early-look-at-kimi-k3-generations-from-moonshot-ai-on-arena/ Last updated: 2026-07-16T11:17:03.000Z Moonshot AI appears to be on the verge of shipping Kimi K3, the successor to its K2 model family, and the launch signals are stacking up fast. Some beta users have reportedly spotted the model in Kimi's own selectors, while a mystery checkpoint called Kivine surfaced on LM Arena, fitting the anonymized pattern labs use to test upcoming models on the platform. Moonshot leaned into the moment itself, posting a short teaser video built around the number three, and a recharge-campaign page referencing a K3 launch briefly appeared on its developer platform before being pulled, pointing at a mid-July window that now looks like days away at most. > [pic.twitter.com/vZrE9vCrU4](https://t.co/vZrE9vCrU4?ref=testingcatalog.com) > > — Kimi.ai (@Kimi\_Moonshot) [July 15, 2026](https://x.com/Kimi%5FMoonshot/status/2077521842080817296?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [Rumors](https://x.com/zijing%5Fwu/status/2077699476194771064?ref=testingcatalog.com) describe a Mixture-of-Experts model in the multi-trillion-parameter range with a 1-million-token context window, though nothing has been confirmed by a model card yet. Early Arena impressions place Kivine near frontier territory. In our own head-to-head against Claude Fable 5 on a universe-simulation prompt, Fable finished faster and produced sturdier UX components, while K3's build was more elaborate and visually rich, at one point spinning the camera into a first-person view around a selected planet. Testers still report weak spots, including long runtimes on harder agent tasks, but the gap to the top models looks unusually narrow for an unreleased checkpoint. > Kivine, a new model on [@arena](https://x.com/arena?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), is potentially an upcoming Kimi K3 model. > > I got lucky and caught a comparison between Fable 5 and Kimi K3 in the Universe simulation prompt. > > \> Fable 5 finished faster, and most UX components were more robust and easy to use. > > \> [@Kimi\_Moonshot](https://x.com/Kimi%5FMoonshot?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) K3… [https://t.co/5KeH4uNn6T](https://t.co/5KeH4uNn6T?ref=testingcatalog.com) [pic.twitter.com/CAP96E6GkK](https://t.co/CAP96E6GkK?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 15, 2026](https://x.com/testingcatalog/status/2077408151880552477?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > Kimi K3 vs GPT-5.6 Sol > > Difference in taste is so stark , like if i swap kimi k 3 name with fable 5 people will trust it > > but this is not only about visuals , its also about function , you can see both achieve same result but the way to achieve is different kimi is more… [pic.twitter.com/zpirIariVT](https://t.co/zpirIariVT?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [July 16, 2026](https://x.com/chetaslua/status/2077701096924229744?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > After trying a few generations with Kimi K3 by [@KimiDevs](https://x.com/KimiDevs?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) . I think one of its biggest strengths might be building interactive 3D experiences. > > The level of detail, polish, and overall quality is honestly wild. [pic.twitter.com/Sb9GTvzI6f](https://t.co/Sb9GTvzI6f?ref=testingcatalog.com) > > — Noctus (@noctus91) [July 15, 2026](https://x.com/noctus91/status/2077455477131419759?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > Holy Shit Kimi 🤯 > > the more i test it the more i am thinking , opensource and china is not that behind from current gen SOTAs > > this is better output than upcoming opus 5 same prompts and different result [https://t.co/hGIK0pDVV5](https://t.co/hGIK0pDVV5?ref=testingcatalog.com) [pic.twitter.com/0Tn9e5ZZkX](https://t.co/0Tn9e5ZZkX?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [July 15, 2026](https://x.com/chetaslua/status/2077421558155690459?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > One of the first Kimi K3 outputs for y'all, tried it on frontend 👀 > > First impressions is that it's VERY slow, even slower than Fable. This took 35 minutes to finish. > > However, this is one of the best outputs I've ever seen from this prompt, better than many frontier models. [pic.twitter.com/H8FjJwGd8v](https://t.co/H8FjJwGd8v?ref=testingcatalog.com) > > — Lentils (@Lentils80) [July 15, 2026](https://x.com/Lentils80/status/2077387333154857151?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > Kimi K3 vs GPT 5.6 Sol (extra high) > > Voxel Death Star trench run > > I do want to point out that this was supposed to be a static world, but when I gave this prompt to Kimi K3, it said: > > “Since I can't output raster images, I've done something better: a real-time, fully animated… [pic.twitter.com/v3n5oubuwE](https://t.co/v3n5oubuwE?ref=testingcatalog.com) > > — Chris (@ChrissGPT) [July 16, 2026](https://x.com/ChrissGPT/status/2077608230440640990?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > kimi k3 just beat claude fable > > Sakura bonsai test: > > K3 nailed the overall bonsai shape with a realistic twisted trunk, layered canopy.......a design that closely followed the prompt > > Impressed with the detailing [@Kimi\_Moonshot](https://x.com/Kimi%5FMoonshot?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [pic.twitter.com/uhkQjRlcat](https://t.co/uhkQjRlcat?ref=testingcatalog.com) > > — Synopsis (@abhinavflac) [July 15, 2026](https://x.com/abhinavflac/status/2077458858595987559?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The stakes are clear. [Moonshot](https://www.testingcatalog.com/tag/kimi/) shipped the K2 line as open weights, and K2.6 was briefly the most popular open model before another model overtook it, so K3 reads as a bid to reclaim that position, backed by a funding round earmarked for its training. If the weights follow precedent, developers and self-hosters would gain a frontier-scale open option aimed at long-context agent and codebase work. An open release is expected rather than guaranteed and could trail the API debut, but with the teaser out and the date already leaked once, the wait should be short. ### Raft 1.0 puts AI agents in Team Mode URL: https://www.testingcatalog.com/raft-1-0-puts-ai-agents-in-team-mode/ Last updated: 2026-07-16T08:26:52.000Z Raft 1.0 is now live, launched by [@istdrc](https://x.com/istdrc?ref=testingcatalog.com), who built Kimi CLI at Moonshot AI, marking the full launch of a multi-agent collaboration platform built around what the company calls team mode. Instead of juggling terminals, sessions, and separate agent windows, users get one shared workspace of channels, threads, tasks, and mentions where humans and AI agents work side by side, and where the work stays attached to the conversation it came from. > Hi, I'm RC. I built Kimi CLI at Moonshot last year, and back in 2015, bots that lived in group chats. For the past four months, I've been building Raft in public. > > Today I'm launching Raft 1.0. > > Right now, working with agents means juggling terminals, sessions, and skills. The… [pic.twitter.com/rkJbsHrwB8](https://t.co/rkJbsHrwB8?ref=testingcatalog.com) > > — stdrc (@istdrc) [July 15, 2026](https://x.com/istdrc/status/2077376131628707907?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Each agent on Raft runs as a persistent process with its own identity, memory, and expertise, on whichever runtime fits the job, including Claude, Codex, and Hermes. Agents claim tasks, run in parallel, hand work to one another, and review each other's output in shared threads, so what one agent figures out, the next one builds on. A separate review agent can catch issues in another agent's code before a human ever sees it, which the company frames as independent agents covering independent blind spots. Execution happens on the user's own hardware through a lightweight daemon, keeping compute, code, and data under the user's control. 0:00 /1:00 1× The 1.0 release rounds out a public beta that added external agent support starting with Hermes, agent-created channels, joint channels across servers, a Login with Raft option alongside an app marketplace, message search, and attachment previews with comments. According to the company, more than 20,000 agent-native builders and teams started building on Raft during the beta, averaging four agents per human, with power users running over sixty. Raft is developed by [Botiverse](https://botiverse.ai/en?ref=testingcatalog.com), founded by Richard, known as RC, who previously built Kimi CLI at Moonshot AI. The team says it runs 99 percent of the company inside Raft itself, with more than ten humans and over one hundred named agents claiming tasks, reviewing each other's work, and retaining context week to week, including for this launch. Nous Research's Hermes Agent has officially partnered with Raft to run as an external agent within its workspaces, and testimonials on the company's site come from founders and engineers at TiDB, Inferact, and BeFreed AI, among others. SPONSORED Start testing with Raft [Learn more ](https://raft.build/?ref=testingcatalog.com) Raft is now available at raft.build, with a free tier that includes channels, tasks, agents on local machines, and 30 days of message history. A Pro plan runs at 8.80 dollars per seat per month, billed annually, with each human counting as one seat and each agent as a tenth of a seat, while an enterprise tier with private deployment and SSO is listed as coming soon. Users bring the AI subscriptions they already pay for, such as Claude or Codex, so the platform layers team coordination on top of existing model access rather than reselling it. ### Thinking Machines debuts open-weight Inkling AI model URL: https://www.testingcatalog.com/thinking-machines-debuts-open-weight-inkling-ai-model/ Last updated: 2026-07-15T22:09:17.000Z Thinking Machines Lab has released Inkling, its first open-weights foundation model, giving developers and companies full access to customize and deploy it. Inkling is available now for fine-tuning through Tinker, while its complete weights can be downloaded as original and NVFP4 checkpoints for NVIDIA Blackwell systems. Inkling is a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active during inference. It supports context windows of up to one million tokens and was pretrained from scratch on 45 trillion tokens spanning text, images, audio, and video. The model processes text, image, and audio inputs natively, with capabilities covering reasoning, coding, tool use, instruction following, visual analysis, speech transcription, and long-form audio understanding. > [https://t.co/FsLr1YySMI](https://t.co/FsLr1YySMI?ref=testingcatalog.com) > > — Hugging Face (@huggingface) [July 15, 2026](https://x.com/huggingface/status/2077477643767681100?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Developers can control Inkling’s thinking effort between 0.2 and 0.99 to balance output quality, latency, and token consumption. Thinking Machines says Inkling can match Nemotron 3 Ultra on Terminal Bench 2.1 while using roughly one-third as many generated tokens. The company positions it as a broad base for custom models rather than the highest-scoring general-purpose model overall. Inkling scored 77.6% on SWE-bench Verified, 97.1% on AIME 2026, 87.2% on GPQA Diamond, and 73.5% on MMMU Pro. It can operate inside coding-agent harnesses, use changing tool schemas, and produce applications, structured artifacts, and longer projects through repeated refinement. Tinker offers Inkling with 64K and 256K context options and a temporary 50% discount. The Inkling Playground provides a chat interface with integrated agentic web search, free for a limited period. API access is also available through Together, Fireworks, Modal, Databricks, and Baseten, with inference support across SGLang, vLLM, TokenSpeed, llama.cpp, and Hugging Face Transformers. Thinking Machines is also previewing Inkling-Small, a 276-billion-parameter Mixture-of-Experts model with 12 billion active parameters. It approaches or surpasses Inkling on several reasoning, instruction-following, vision, and audio tests while targeting lower-cost and lower-latency workloads. Its full weights will follow after testing is completed. Thinking Machines Lab is an AI research and product company founded by former OpenAI CTO Mira Murati. Inkling extends the company’s Tinker customization platform and is intended to serve as the reasoning layer behind its previously previewed real-time voice and vision systems. The release marks the start of a planned family of customizable models across multiple sizes. [Source](https://thinkingmachines.ai/news/introducing-inkling/?ref=testingcatalog.com) ### OpenAI prepares Codex Micro keypad to control AI Agents URL: https://www.testingcatalog.com/openai-prepares-codex-micro-keypad-to-control-ai-agents/ Last updated: 2026-07-15T16:11:54.000Z Update: Officially announced on the [OpenAI blog](https://openai.com/supply/co-lab/work-louder/?ref=testingcatalog.com). OpenAI's first hardware product arrives today, and new details reveal how deeply the Codex Micro is integrated with the company's coding agent. Built with Work Louder, the keypad expands on the teaser OpenAI shared in late June, which confirmed a July 15 launch and drew comparisons to Work Louder's Creator Micro 2, a $199 macro pad with 13 mechanical keys, a joystick, and a rotary dial. > Your favorite Codex shortcuts are getting an upgrade. > > July 15th. [pic.twitter.com/xZ1ydZyt94](https://t.co/xZ1ydZyt94?ref=testingcatalog.com) > > — OpenAI Developers (@OpenAIDevs) [June 29, 2026](https://x.com/OpenAIDevs/status/2071639953927438440?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) According to an early look at the official render and accompanying description, the standout element is a row of 6 frosted Agent Keys that display the live status of Codex threads through RGB colors: 1. White for idle 2. Blue for thinking 3. Green for complete 4. Amber when input is required 5. Red for errors 6. Off when no agent is running A single tap focuses an agent in the background, while a double tap brings the Codex window to the front. This turns the keypad into a physical dashboard for parallel agent runs, addressing a real pain point for developers juggling several Codex threads who currently rely on window switching to check progress. > [https://t.co/vm3oPcOAeD](https://t.co/vm3oPcOAeD?ref=testingcatalog.com) [pic.twitter.com/JRMgoqVJms](https://t.co/JRMgoqVJms?ref=testingcatalog.com) > > — Tibor Blaho (@btibor91) [July 15, 2026](https://x.com/btibor91/status/2077379861031571481?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Command Keys hold common Codex actions, and long-pressing the dial inside Codex opens a configuration page where those keys, the dial, and the joystick can be remapped. Beyond Codex, Work Louder's Input software assigns shortcuts for any app across 6 programmable layers, toggled through a bottom-left touch sensor with 3 layer LEDs or switched automatically with AppSense, which detects the app in focus after 5 seconds. The launch fits [OpenAI's](https://www.testingcatalog.com/tag/chatgpt/) push to make Codex a daily driver for professional developers, a platform that reportedly passed 5 million weekly active users in June. It also lands as agent-driven coding shifts work from typing to supervising, a mode where dedicated status hardware makes sense. The device is separate from the consumer product OpenAI is developing with Jony Ive, which is still months away. Pricing has not been confirmed, though the Creator Micro 2's $199 tag offers a likely reference point. ### Telegram adds rich text editor, communities, more GIFs URL: https://www.testingcatalog.com/telegram-adds-rich-text-editor-communities-more-gifs/ Last updated: 2026-07-14T23:58:16.000Z Telegram has rolled out a major update featuring a rich text editor, Communities, private bot messages inside groups, and a rebuilt GIF search covering more than 350 million animations. The new rich text editor allows users to create long-form posts with headings, tables, lists, quotes, code blocks, formulas, photos, and videos placed directly between paragraphs. Messages can contain up to 32,768 characters, bringing article-style publishing directly into Telegram chats and channels. The visual editor is optimized for mobile and desktop, while Cocoon-powered AI tools can generate text, formulas, and tables with a focus on privacy. The editor is currently limited to Telegram Premium subscribers. 0:00 /0:14 1× Telegram Telegram Communities introduce a new way to organize multiple groups, channels, and bots around a single topic. Members can discover and join connected chats without separate invitation links. A community can appear as individual chats or as one expandable item in the chat list. Community administrators can decide whether connected chats are visible to everyone or hidden from people who are not already members. Communities are collaborative by default, meaning members can add chats. Administrators can restrict this behavior, turning new additions into suggestions that require approval. 0:00 /0:18 1× Bots inside groups can now send ephemeral messages visible only to a specific user. These private responses can include AI summaries, greetings, errors, confirmations, or menus with buttons, without cluttering the public conversation with messages intended for one person. Bots can send them automatically or after a user selects a command marked with a dedicated icon. 💡 Join [TestingCatalog channel](https://t.me/testingcatalog?ref=testingcatalog.com) on Telegram too! Telegram has also replaced its GIF search system with an internally managed index containing more than 350 million publicly available GIFs across 36 languages. The index was built using open-source AI models and works without third-party search providers, keeping search queries within Telegram’s infrastructure. The company says the new system also returns results faster than its previous GIF search. The update pushes Telegram further beyond conventional messaging by combining long-form publishing, community management, private bot responses, and media discovery inside the same platform. Telegram operates cloud-based apps across Android, iOS, desktop, macOS, and the web, allowing chats and content to remain synchronized across devices. [Source](https://telegram.org/blog/communities-editor-invisible-messages?ref=testingcatalog.com) ### Exclusive: Early 30-second AI videos generated by Seedance 2.5 URL: https://www.testingcatalog.com/exclusive-early-30-second-ai-videos-generated-by-seedance-2-5/ Last updated: 2026-07-11T23:05:44.000Z ByteDance is closing in on the release of Seedance 2.5, the next step in its video model line and the first to generate [native 30-second clips](https://www.testingcatalog.com/bytedance-set-to-launch-seedance-2-5-with-3-minute-ai-video-output/) in a single pass. The launch was originally penciled in around July 9 but slipped, reportedly amid ongoing talks with rights holders, and the API release is now targeted for July 16\. Given how previous Seedance rollouts unfolded, the model may well surface earlier in some places: references to it have been multiplying across partner apps that bundle Seedance models into their video generation and editing toolkits, and pricing details have already appeared on Jimeng in China. > Ultra-Long Video (Beta), powered by SeeDance 2.5 has been spotted on Jimeng. > > This mode has been announced earlier and will allow users to generate 180 second long videos! > > h/t [@MarsForTech](https://x.com/MarsForTech?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [https://t.co/5NkvdsewQD](https://t.co/5NkvdsewQD?ref=testingcatalog.com) [pic.twitter.com/Z2tOQ6E3Bg](https://t.co/Z2tOQ6E3Bg?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 9, 2026](https://x.com/testingcatalog/status/2075160295966642333?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) We have had early access to the model, and the first results are hard to argue with. Every clip we generated was a one-shot, 30-second video, and the consistency of characters, lighting, and pacing across that duration is unlike anything else we have tested. Small details occasionally call for a regeneration, but the output holds together remarkably well. The model, unveiled at ByteDance's FORCE conference in June, also accepts up to 50 reference inputs spanning images, video, and audio, and supports region-level edits that change part of a frame without redoing the whole clip. > BYTEDANCE 🔥: Some exclusive Seedance 2.5 Pro samples have arrived. > > “Cyberpunk hacker robot working in front of many monitors” test. > > 30 seconds one shot generation 👀 [https://t.co/HWC6H9B7Xo](https://t.co/HWC6H9B7Xo?ref=testingcatalog.com) [pic.twitter.com/abTrfa9now](https://t.co/abTrfa9now?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 11, 2026](https://x.com/testingcatalog/status/2076043332472463471?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The rollout is expected to run through Jimeng and Dreamina first, then CapCut and the Volcano Engine API, putting the model in front of advertisers, short-form creators, and ecommerce teams for whom a 30-second spot in one pass removes the stitching step entirely. [Join Dev Mode ](https://discord.com/invite/devmode?ref=testingcatalog.com) The stakes are as competitive as they are technical. Seedance 2.0 topped independent video arenas earlier this year, even as Hollywood cease-and-desist letters slowed its global rollout, and ByteDance has since layered in filters and previewed a copyright-licensing platform. [Google's Gemini Omni](https://www.testingcatalog.com/google-rolls-out-gemini-omni-ai-for-video-generation-and-editing/) came close in video editing, but Omni Flash tops out at around 10 seconds per clip, so a coherent one-shot half-minute would reset the gap once again. Mid-July should tell. [Source](https://x.com/MarsForTech?ref=testingcatalog.com) ### OpenAI AMA: Reddit summary about new ChatGPT and GPT-5.6 URL: https://www.testingcatalog.com/openai-ama-reddit-summary-about-new-chatgpt-and-gpt-5-6/ Last updated: 2026-07-11T23:11:15.000Z OpenAI’s Codex team has disclosed new operating guidance for GPT-5.6 and Codex during a [Reddit AMA](https://www.reddit.com/r/codex/comments/1us9ty9/ama%5Fwith%5Fopenais%5Fcodex%5Fteam/?ref=testingcatalog.com) held on July 10, alongside confirmation that Codex now reaches more than five million weekly users. That figure has doubled in three months, while the team says it shipped 150 features and product changes during the same period. > Summary of Reddit AMA about "GPT-5.6 and Codex in ChatGPT" with OpenAI's Codex team on 2026-07-10 > > (opened with the stat that more than 5 million people use Codex every week, twice as many as three months ago, with 150 features and improvements shipped in that period) > > Model… [https://t.co/FfFeVibJpS](https://t.co/FfFeVibJpS?ref=testingcatalog.com) [pic.twitter.com/7bv7fid9hO](https://t.co/7bv7fid9hO?ref=testingcatalog.com) > > — Tibor Blaho (@btibor91) [July 10, 2026](https://x.com/btibor91/status/2075670159264461255?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For most coding work, **OpenAI recommends** [**GPT-5.6**](https://www.testingcatalog.com/openai-launches-gpt-5-6-sol-terra-and-luna-on-apps-and-api/) **Sol at Medium reasoning.** Sol Ultra is positioned for migrations, security-sensitive changes, production incidents, and other tasks where mistakes are costly. Terra is intended for faster or usage-conscious work, while Luna is suited to lightweight operations and subagents. There is no automatic model router yet, although the new reasoning slider can fall back to Terra at its lowest setting. The team said Sol is its strongest option for interface work, especially when reference images are supplied. Sol Medium is also described as faster than GPT-5.5 for many tasks, while Fast mode runs at roughly 1.5 times the normal speed. OpenAI is preparing a Cerebras-backed configuration targeting about 750 tokens per second, but has made no commitment to a one-million-token context window for Sol. Codex users were also advised to **give agents bounded goals, require tests after each attempt**, and use the “/goal” command for longer investigations. OpenAI acknowledged that models can abandon patches too quickly when results are imperfect. Greater persistence and lower code complexity are now listed as areas for further work. Agentic usage is counted by the feature being used rather than the device or interface. Codex in the desktop app, CLI, IDE, web, and mobile draws from the same agentic allowance as [ChatGPT Work](https://www.testingcatalog.com/openai-launches-chatgpt-work-for-pro-enterprise-and-edu-plans/), while standard ChatGPT conversations use a separate allocation. OpenAI said task cost can vary sharply depending on repository size, reasoning depth, and execution time, and that it is working on clearer usage reporting. The company also addressed criticism of the merged ChatGPT desktop app. ChatGPT Classic and the new app can run side by side for now, while fixes are being applied to browser control, stuck threads, connection failures, packaging, and resource usage. Windows parity is receiving additional attention, and a Linux desktop app is under development, with no public timeline. > Notes on GPT-5.6, which includes some interesting new additions to the API (programmatic tool calling and multi-agent in particular) - plus 18 pelicans for the 6 reasoning levels and 3 new models:[https://t.co/pZOrG6QxTL](https://t.co/pZOrG6QxTL?ref=testingcatalog.com) > > — Simon Willison (@simonw) [July 9, 2026](https://x.com/simonw/status/2075306164993315192?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Early reactions remain mixed. Simon Willison said Sol Medium may become his default for coding, while also describing the selection of models and reasoning as confusing. [Ethan Mollick](https://x.com/emollick?ref=testingcatalog.com) questioned whether ChatGPT Work has a clear role beside Codex, reflecting continued uncertainty about how OpenAI is separating its general-purpose agent from its developer-focused workspace. [Source](https://x.com/btibor91/status/2075670159264461255?ref=testingcatalog.com) ### Cursor 3.11 adds side chats and agent transcript search URL: https://www.testingcatalog.com/cursor-3-11-adds-side-chats-and-agent-transcript-search/ Last updated: 2026-07-11T22:32:21.000Z Cursor has released version 3.11, introducing Side Chats, a new feature that allows developers to open parallel agent conversations without interrupting or redirecting their primary coding tasks. Developers can initiate a side chat using commands like `/side`, `/btw`, or by clicking the plus button in the chat panel. Each side chat receives context from the main thread, remains accessible for future follow-ups, and can be referenced with an @-mention to integrate its findings back into the original conversation. > Introducing side chats, a new way to ask questions and explore ideas without interrupting your main conversation. > > Each side chat is a durable agent conversation you can @-mention to bring context back into the main thread. [pic.twitter.com/bc6YlhJWWH](https://t.co/bc6YlhJWWH?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [July 10, 2026](https://x.com/cursor%5Fai/status/2075686268113916023?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Side chats are designed to read, search, and answer queries, enabling developers to explore alternatives, clarify implementation details, or review decisions while the main agent continues its work. Unlike starting an unrelated chat, this new system maintains the connection between the primary task and each tangent, transforming Cursor’s agent interface into a workspace for multiple concurrent reasoning paths. Cursor 3.11 also introduces a full conversation search feature. The Agents Window now allows users to search transcript content using Cmd+K, rather than relying solely on conversation names and pull request numbers. According to Cursor, this feature uses a local index that can handle thousands of conversations. Additionally, Cmd+F enables searches within the current transcript, complete with match navigation and counters for lengthy sessions. The release reorganizes project and repository selection around three execution locations: This Computer, Cloud, and Remote Machines. Users can create projects, connect to GitHub, GitLab, or Azure DevOps, select multiple repositories, and choose branches without leaving the picker. Cloud agents also gain hooks for prompts, responses, thoughts, subagents, compaction, stopping, and turn completion. These controls can be used to inspect agent activity, apply policies, or create self-correcting workflows. Cursor is developed by Anysphere, an applied research lab focused on the future of programming. Version 3.11 advances Cursor from being a single coding assistant to a coordinated agent workspace. This allows developers to run a primary task, investigate related questions in parallel, retrieve prior work, and control cloud agent behavior through programmable lifecycle hooks. [Source](https://cursor.com/changelog/side-chat?ref=testingcatalog.com) ### Drafted released a major V2 upgrade to its AI home design tool URL: https://www.testingcatalog.com/drafted-released-a-major-v2-upgrade-to-its-ai-home-design-tool/ Last updated: 2026-07-10T17:21:29.000Z Drafted has released the second version of its AI home design platform and kept the tool free to use while rebuilding the model that generates the plans. The product turns structured inputs into complete residential layouts, taking a room list, square footage, a lot boundary, and a footprint shape and returning a full floor plan with a matching exterior that can be reviewed in 2D and 3D. Finished plans are exported as CAD, PDF, and BIM files for the rest of the pre-construction process. The center of V2 is a new in-house model that Drafted reports runs 60 percent faster than the previous generation, achieves 5 times higher plan-adherence accuracy, produces cleaner hallways and circulation, and places rooms 33 percent better within irregular shapes. V2 also adds controls that change how a plan gets built. Room Placement lets someone drop a room at an exact size and location, then generates the rest of the house around that fixed point. Regeneration runs only on selected areas, so clicking a part of the layout returns fresh options for that section without redrawing the whole plan. Roof editing brings a switch between hipped and gable styles, and a live 3D view updates in real time as the porch, house shape, or roofline changes. ![Drafted](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Drafted-Design-your-dream-house-plan-with-AI-07-10-2026_01_58_PM.jpg) Two more capabilities extend the canvas: a user can sketch a footprint and have rooms generated to fill that outline, and any home created on the site can be remixed into the starting point for a new design. Drafted plans to announce additional features over the next two weeks. The tool centers on single-family homes and sits at the schematic stage, where early concepts are usually drawn, discarded, and redrawn before a direction is chosen. SPONSORED Start testing with Drafted AI [Learn more ](https://drafted.ai/?ref=testingcatalog.com) Drafted is a San Francisco company founded in 2025 by Nicholas Donahue, part of Y Combinator's recent batch, with a team of around nine. It raised a $16 million seed round from backers including Y Combinator, Buckley Ventures, Pinterest cofounder Ben Silbermann, and Ryan Tedder. Donahue previously founded Atmos, a custom home building company, and frames Drafted as a more AI-native run at the same problem: making early home design cheaper and faster to explore. The company trained its own model on real house plans from homes that were built and had permitting cleared, an approach it credits with both accuracy and low running costs. ![Drafted](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Drafted-Design-your-dream-house-plan-with-AI-07-10-2026_02_01_PM.jpg) Drafted reports that more than 120,000 people generated over 325,000 designs in a recent month, across a user base that spans homeowners and home buyers, builders, drafters and architects, developers, agents, and interior designers. The stated goal is to make exploring the built environment as approachable as editing a website, and to move a project from an idea to a decision without supplanting the professionals who carry a plan into formal design. ### NoimosAI launches website builder that automates SEO URL: https://www.testingcatalog.com/noimosai-launches-website-builder-that-automates-seo/ Last updated: 2026-07-10T16:03:43.000Z NoimosAI has released its Website Builder, a feature it describes as the Self-Improving Web, giving users a way to build, publish, and drive traffic to a website without writing code or handling manual configuration. The framing is that a site should not only exist but also continue to attract visitors, generate content, and refine its performance around the clock, rather than sit idle after launch. > Introducing NoimosAI Website Builder: The Self-Improving Web. > > Now you can build, launch, and drive traffic to a beautiful website 24/7 without writing code, dealing with complex setups, or creating content manually. > > Your website shouldn't just exist—it should get traffic. [pic.twitter.com/ziWOySsJ3G](https://t.co/ziWOySsJ3G?ref=testingcatalog.com) > > — NoimosAI (@noimos\_ai) [July 10, 2026](https://x.com/noimos%5Fai/status/2075610850636001745?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The build starts from a chat. A user describes the business or the site they have in mind, and NoimosAI generates an SEO and GEO-optimized site in real time, automatically producing an llms.txt file and JSON-LD structured data so that AI models can surface and recommend the business in generated answers. Teams that already have code can bring it over via a direct GitHub repository import. Publishing runs without DNS work: domains can be purchased, and the site can be pushed live from within the platform, removing the configuration step that usually sits between a finished build and a public URL. ![NoimosAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/image-1.png) The part NoimosAI leans on hardest is what happens after launch. The system reads performance data from Google Search Console and Google Analytics, incorporates recommendations from multiple AI models, and rewrites its own code and HTML to address performance bottlenecks over time. Alongside that, NoimosAI agents autonomously generate SEO, GEO, and social content to keep traffic flowing, so the site is treated as an operating surface that is maintained continuously rather than a one-off deliverable. The intended audience is founders, freelancers, creators, marketers, and small-business owners who want a working site and a steady stream of visitors without a dedicated developer or a separate content operation. SPONSORED Start testing NoimosAI! [Take me there! ](https://noimosai.com/en?ref=testingcatalog.com) The Website Builder sits inside the broader NoimosAI platform, billed as the world's first all-in-one autonomous AI marketing team that runs the full cycle from strategy planning through execution and continuous improvement. Built by AGOS LABS, the platform connects a brand's apps, websites, and data sources, then plans and runs growth work across SEO, GEO, social, and outreach, with finished outputs routed back for review and approval. It builds a knowledge base and memory from a connected site or set of accounts, learns a brand's style and real-time metrics such as search traffic and conversion rates, and adjusts its personalization and output from feedback and results. The Website Builder extends that same loop to the site layer itself, folding creation, publishing, and ongoing optimization into the autonomous system the company has been building toward since its platform launch earlier this year. ### OpenAI launches ChatGPT Work for Pro, Enterprise, and Edu plans URL: https://www.testingcatalog.com/openai-launches-chatgpt-work-for-pro-enterprise-and-edu-plans/ Last updated: 2026-07-09T21:54:11.000Z OpenAI is transforming ChatGPT into a comprehensive work agent with the introduction of ChatGPT Work, a new mode designed to perform actions across connected apps, files, web tools, desktop software, and recurring workflows. This release extends ChatGPT's capabilities beyond merely answering prompts to executing longer projects, supporting the creation of sheets, slide decks, documents, dashboards, web apps, and team-ready materials from a single work request. [ChatGPT Work](https://www.testingcatalog.com/openai-set-to-launch-chatgpt-work-upgrades-today/) is initially available on web and mobile for Pro, Enterprise, and Edu users, with Plus and Business access expected in the coming days. The updated ChatGPT desktop app is now globally available on Windows and Mac, integrating Chat, Work, and Codex into one app across all plans, including Free. OpenAI notes that the older desktop app will transition to ChatGPT Classic, while the Codex app is being integrated into the new ChatGPT desktop app. > Introducing ChatGPT Work, a new agent in ChatGPT powered by Codex and GPT-5.6. > > It can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work. > > It’s a whole new way to get work done. [pic.twitter.com/uGbvjU1LsV](https://t.co/uGbvjU1LsV?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [July 9, 2026](https://x.com/OpenAI/status/2075274271845404744?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The new agent is powered by GPT-5.6, OpenAI’s latest frontier model, and incorporates Codex technology for multi-step execution. ChatGPT Work can decompose larger goals into smaller tasks, continue working for extended periods, seek guidance when necessary, and request approval before executing sensitive actions. It can utilize plugins for Slack, Microsoft Teams, Gmail, Google Drive, SharePoint, Salesforce, calendars, project trackers, CRMs, and other workplace systems. Users can also invoke a specific app by typing “@” followed by the plugin name. OpenAI is also launching Sites in public beta, enabling users to convert work materials into shareable interactive sites or web apps. This feature can encompass internal portals, dashboards, launch calendars, project trackers, prototypes, and interactive reports. ChatGPT can test these sites within the app and update them when source information changes. > ChatGPT Sites is now available to Plus, Pro, Business and Enterprise. On web, mobile, and desktop. > > I turned an idea into a personal to-do app with the Sites plugin in the new ChatGPT. > > Here’s how I built it. [pic.twitter.com/HiNDUxZ32O](https://t.co/HiNDUxZ32O?ref=testingcatalog.com) > > — pranav (@prd\_008) [July 9, 2026](https://x.com/prd%5F008/status/2075326663852953929?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) On desktop, ChatGPT now includes a built-in browser and Computer Use capabilities. This allows it to gather web context, operate through browser-based tools, open and refine Google Workspace or Microsoft 365 files, click, type, and move files across local apps. OpenAI is also updating its Chrome extension to position ChatGPT in the Chrome sidebar and initiating the shutdown of its standalone Atlas browser. > The new ChatGPT Work comes with a new Computer Use experience. > > It's faster and introduces picture-in-picture, so you can keep an eye on Computer Use while it works! [pic.twitter.com/uSqbZdAcYv](https://t.co/uSqbZdAcYv?ref=testingcatalog.com) > > — Ari Weinstein (@AriX) [July 9, 2026](https://x.com/AriX/status/2075282339782095163?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The release is aimed at knowledge workers, enterprise teams, students, developers, and business users who require [ChatGPT](https://www.testingcatalog.com/tag/chatgpt/) to manage workflows rather than isolated responses. OpenAI provides examples across sales, marketing, finance, business operations, analytics, and engineering, including account planning, launch checks, month-end finance work, market research, and recurring executive reporting. Early testers identified by OpenAI include Zapier, RingCentral, Virgin Atlantic, and NVIDIA. Reported use cases involve reviewing thousands of sales leads, checking launch readiness across Jira and go-to-market plans, comparing airline passenger experiences for a five-year strategy, and automating event preparation around NVIDIA GTC. For organizations, ChatGPT Work is linked to ChatGPT Enterprise controls. Admins can manage access to plugins, browser use, connected tools, network access, sensitive actions, spend controls, and usage limits. OpenAI states that auto-review employs advanced models to verify critical connected-tool and API actions before they occur, and that internal red-team testing prevents attempts to extract protected data. [Source](https://openai.com/index/chatgpt-for-your-most-ambitious-work/?ref=testingcatalog.com) ### OpenAI launches GPT-5.6 Sol, Terra, and Luna on apps and API URL: https://www.testingcatalog.com/openai-launches-gpt-5-6-sol-terra-and-luna-on-apps-and-api/ Last updated: 2026-07-09T21:04:51.000Z OpenAI has released GPT-5.6, its new frontier model family for ChatGPT, ChatGPT Work, Codex, and the OpenAI API. The rollout starts globally today and is set to reach full availability over the next 24 hours. The family includes three tiers: Sol as the flagship model, Terra as the lower-cost everyday work option, and Luna as the fastest and most affordable model. The launch moves GPT-5.6 from a limited preview into general availability. OpenAI is positioning the series around higher intelligence per token, lower estimated cost for complex work, and stronger agentic performance across coding, knowledge work, cybersecurity, science, design, and internal research workflows. Sol is the top-tier model, while Terra and Luna are intended to make the same generation available at lower cost and latency. > Sol, Terra, and Luna, our GPT‑5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API. [pic.twitter.com/Qri7GdtYs3](https://t.co/Qri7GdtYs3?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [July 9, 2026](https://x.com/OpenAI/status/2075271421149020426?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) A major addition is the new ultra setting, which coordinates multiple agents across parallel workstreams for demanding tasks. OpenAI says ultra uses four agents by default, trading higher token use for stronger results and faster completion on complex work. Developers can build similar workflows through the multi-agent beta in the Responses API, while Programmatic Tool Calling lets GPT-5.6 write and run in-memory JavaScript to coordinate tools, call them in parallel, use loops and conditions, and process intermediate results before returning an answer. GPT-5.6 is also aimed at professional artifact generation. OpenAI says the model can create editable presentations, documents, spreadsheets, interfaces, visual explanations, and frontend prototypes with stronger layout judgment and closer adherence to reference files. In ChatGPT Work, the model is designed to handle source material from documents and connected work apps, then convert it into shareable outputs. This puts GPT-5.6 directly into OpenAI’s broader push to make ChatGPT a work-execution environment rather than just a conversational assistant. > GPT-5.6 System Card[https://t.co/9wCeJAe7Fx](https://t.co/9wCeJAe7Fx?ref=testingcatalog.com)… [pic.twitter.com/5uQfbjFue4](https://t.co/5uQfbjFue4?ref=testingcatalog.com) > > — Andrew Curran (@AndrewCurran\_) [July 9, 2026](https://x.com/AndrewCurran%5F/status/2075266557895459266?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For developers, GPT-5.6 arrives with Sol, Terra, and Luna in the API. Pricing is set at $5 input and $30 output per 1 million tokens for Sol, $2.50 input and $15 output per 1 million tokens for Terra, and $1 input and $6 output per 1 million tokens for Luna. The release also adds more predictable prompt caching, explicit cache breakpoints, and a 30-minute minimum cache life. Cache writes are billed at 1.25x the uncached input rate, while cache reads retain a 90% discount on the cached input rate. > On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol sets a new state of the art at 80.0—2.8 points above Claude Fable 5—while using less than half the output tokens, taking less than half the time, and costing about one-third less. [pic.twitter.com/H5o0qJKUOL](https://t.co/H5o0qJKUOL?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [July 9, 2026](https://x.com/OpenAI/status/2075271425548795909?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Access depends on product and plan. In ChatGPT, Plus, Pro, Business, and Enterprise users get GPT-5.6 Sol through medium- and higher-effort settings, while Pro and Enterprise users can select Sol Pro for complex tasks. In ChatGPT Work and Codex, Free and Go users get Terra, while Plus, Pro, Business, and Enterprise users can choose Sol, Terra, and Luna with effort controls. The max setting is available to all users with access to GPT-5.6 in ChatGPT Work and Codex. Ultra is available to Pro and Enterprise users in ChatGPT Work and to Plus and higher plans in Codex. OpenAI says GPT-5.6 Sol sets new or near-frontier results across several evaluations. The company highlights gains in coding with Terminal-Bench 2.1 and DeepSWE; knowledge work with BrowseComp and OSWorld 2.0; cybersecurity with ExploitBench, ExploitGym, and SEC-Bench Pro; and scientific workflows with GeneBench Pro, LifeSciBench, and chemistry-related evaluations. The company also says GPT-5.6 is now used internally by OpenAI researchers for debugging systems, optimizing training, running experiments, and interpreting results. [Source](https://openai.com/index/gpt-5-6/?ref=testingcatalog.com) ### Vellum launches Plugin Hub for personal AI assistants URL: https://www.testingcatalog.com/vellum-launches-plugin-hub-for-personal-ai-assistants/ Last updated: 2026-07-09T18:06:33.000Z Vellum has launched Plugin Hub, an open catalog of installable plugins for its local-first personal AI assistant. This release distinguishes between skills and plugins: while a skill handles a single task such as drafting one email, a plugin provides the assistant with a standing capability it can execute independently. This includes tasks such as connecting to Gmail and reading your inbox, managing a calendar, tracking expenses, syncing notes, or monitoring code repositories. Each plugin is positioned as a complete function the assistant owns, rather than a one-off prompt. 0:00 /0:09 1× Plugins are developed as TypeScript packages with lifecycle hooks, differing from skills, which are instruction bundles. The manifest utilizes the vellumai plugin-api and exposes hooks for init, user-prompt-submit, post-tool-use, post-model-call, and on Shutdown, with state persisted to a local JSONL file. This design is intended to be more accessible than authoring a full MCP server, allowing more people than just infrastructure engineers to build one. For those who prefer not to write code, Vellum offers a plugin builder skill that generates a plugin from a plain-language description. A key feature of this setup is credential handling. When a plugin requires an API key or an OAuth token to access email, banking, or a code repository, that credential remains on the machine and is not sent to the model. This ensures the assistant can be trusted with sensitive integrations. Plugins are open-source and portable, allowing them to be shared among users rather than tied to a single vendor. A single plugin can also operate across different platforms: installed on macOS, the same plugin and assistant can run through Slack and Telegram with shared memory, making a capability built once available wherever the assistant is accessed. The launch catalog includes a variety of use cases: 1. **Marketing Expert:** Bundles positioning, launches, content, brand voice, and reporting into one plugin. 2. **AI Hero Engineer Kit:** Packages engineering workflows around plans, PRDs, vertical-slice issues, and test-driven loops. 3. **Writing Coach:** Runs daily prompts that escalate with a streak and critiques drafts. 4. **Fitness Companion:** Logs meals, workouts, and weight, pulls nutrition data from OpenFoodFacts, and calculates macro targets. Plugins can be installed from the catalog or directly from a GitHub URL, and authors can publish their own to the same catalog. SPONSORED Start testing Vellum! [Take me there! ](https://vellum.ai/?utm%5Fsource=twitter&utm%5Fmedium=social&utm%5Fcampaign=feature-launch-07-09) Vellum is a local-first personal AI assistant that operates as a native macOS app or as a self-hosted service, with an open-source runtime and support for multiple model providers, including Anthropic, OpenAI, Google, and local models via Ollama. The assistant is built around persistent memory, its own identity, and permission-gated action on the user's machine. Plugin Hub extends this foundation into a shared ecosystem, transitioning the assistant from handling one-off tasks to managing ongoing parts of a workflow. The company describes the catalog as the layer that transforms the assistant from a product into a platform. ### Anthropic adds usage reflection dashboard to Claude for all users URL: https://www.testingcatalog.com/anthropic-adds-usage-reflection-dashboard-to-claude-for-all-users/ Last updated: 2026-07-09T15:29:16.000Z Anthropic is introducing a new reflection dashboard to Claude, providing users with a way to review their chatbot usage across recurring topics, task types, and usage patterns. This feature is launching in beta for Free, Pro, and Max users who have Memory turned on, accessible through Settings on Claude for web and the desktop app. Support for cowork conversations is not yet live and is listed as coming soon. The new tool summarizes a user’s Claude activity over the past 1, 3, 6, or 12 months, transforming that history into a usage report. It can display the types of tasks a user frequently brings to Claude, when they use it most, and how that work aligns with Anthropic’s AI Fluency Framework, which includes delegation, description, discernment, and diligence. Anthropic indicates that future updates will provide a view of how much time users have spent using Claude. > Introducing a new way to reflect on how you use Claude. > > Your monthly recap shows when you use Claude most and what you spent that time working on, with options to set quiet hours and nudges to take breaks. Find your dashboard in Settings under Reflect: [https://t.co/8QAn47W5rI](https://t.co/8QAn47W5rI?ref=testingcatalog.com) [pic.twitter.com/WzA3JfONlL](https://t.co/WzA3JfONlL?ref=testingcatalog.com) > > — Claude (@claudeai) [July 9, 2026](https://x.com/claudeai/status/2075225957711929667?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The dashboard also introduces controls focused on self-management rather than on productivity scoring. Users can set quiet hours, schedule break nudges after a chosen amount of Claude use, and respond to reflection prompts about tasks they still want to handle themselves, even when Claude could assist. These reminders are tailored to user preferences and can be dismissed. Privacy is a core aspect of the rollout. Anthropic states that the reflection report does not draw from incognito chats, does not pull underlying files from connected tools, and excludes conversations linked to health integration tools. Sensitive conversations may still appear, but only at a high level. The company assures that the insights remain within the reflection experience and are not used for other purposes. [Anthropic](https://www.testingcatalog.com/tag/claude/), the company behind Claude, describes itself as an AI safety and research company focused on building reliable, interpretable, and steerable AI systems. This launch aligns with that positioning by treating AI usage as something users can inspect and adjust, not just increase. For Claude users, the immediate change is practical: anyone on a supported consumer plan with Memory enabled can generate a report from Settings to determine whether their Claude habits align with their goals. [Source](https://www.anthropic.com/news/reflect-with-claude?ref=testingcatalog.com) ### Meta debuts Muse Spark 1.1 model and opens API for developers URL: https://www.testingcatalog.com/meta-debuts-muse-spark-1-1-model-and-opens-api-for-developers/ Last updated: 2026-07-09T15:10:11.000Z Meta is rolling out Muse Spark 1.1, a new multimodal reasoning model from Meta Superintelligence Labs, alongside the public preview of the Meta Model API. The model is now available in Thinking mode in Meta AI, while developers can start using it via Meta’s new API surface. Meta says the release brings major gains in agentic workflows, computer use, coding, and multimodal understanding compared with the first Muse Spark model. > 1/ muse spark 1.1 is an industry-competitive agentic and coding model. across many agentic evals it rivals gpt-5.5 and opus-4.8. > > available now through the new meta model api and in meta ai. 🧵 [pic.twitter.com/AfaAPDPPxC](https://t.co/AfaAPDPPxC?ref=testingcatalog.com) > > — Alexandr Wang (@alexandr\_wang) [July 9, 2026](https://x.com/alexandr%5Fwang/status/2075218936266998230?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Muse Spark 1.1 is built for long, tool-heavy tasks. It can plan work, call tools, operate across external apps and services, use MCP servers and custom skills, and coordinate parallel subagents. Meta says the model can manage a 1-million-token context window, retrieve information from much earlier in a task, and compact the context so later steps retain critical details. In computer-use scenarios, the model can decide whether to write scripts, click through interfaces, or batch actions depending on the task. > Muse Spark 1.1 excels at multi-app computer-use workflows by maintaining context across extended sessions and intelligently choosing between scripting, direct UI interaction, and batched actions at each step. It’s able to navigate unfamiliar interfaces with minimal human… [pic.twitter.com/qfGueYUj1V](https://t.co/qfGueYUj1V?ref=testingcatalog.com) > > — AI at Meta (@AIatMeta) [July 9, 2026](https://x.com/AIatMeta/status/2075221095880565096?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For coding, Meta positions Muse Spark 1.1 as a large-codebase model for bug fixing, feature implementation, migrations, automated screenshots, debugging, and validation loops. The company says it has trained the model to support agentic coding setups with planning mode, goal conditioning, subagent delegation, and context compaction. The model also adds stronger multimodal workflows, including visual-to-code generation, image and video captioning, and tasks that require the model to inspect visual or audio inputs while acting on a user’s behalf. The new API marks a shift from the earlier Muse Spark rollout, which was limited to Meta AI and a private API preview. Meta’s evaluation report states that Muse Spark 1.1 extends access to external developers via an API that supports tool calling, function calling, and developer prompts. The report also compares it with Muse Spark 1.0 and notes that the new version achieves higher benchmark performance across multiple domains, especially cybersecurity and agentic coding, while employing additional mitigations before deployment. > Meta just released Muse Spark 1.1 and is the new SOTA on MedScribe and TaxEval, taking the top spot from Fable 5 while being 10x cheaper and twice as fast. Meta currently holds the top 2 spots on TaxEval > > It is also the new #1 on Harvey's Legal Agent Bench, dethroning Grok 4.5… [pic.twitter.com/0IjinLA6lq](https://t.co/0IjinLA6lq?ref=testingcatalog.com) > > — Vals AI (@ValsAI) [July 9, 2026](https://x.com/ValsAI/status/2075230620469338210?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Meta says the model was evaluated under its Advanced AI Scaling Framework before release. The evaluation report says that the unmitigated Muse Spark 1.1 reached a high-risk threshold in the chemical and biological, as well as cybersecurity, domains, but Meta applied multi-layered safeguards and assessed residual risk as moderate or lower before launch. The company also reported lower jailbreak attack success than Muse Spark 1.0 on StrongREJECT v2 and lower prompt-injection attack success on AgentDojo. > Meta AI is now powered by Muse Spark 1.1 👀 > > Overall, Meta is one of the companies that can scale their AI to billions of users in a very short amount of time. > > The fact that we saw a 1.1 upgrade coming so fast means only one thing: Meta is back 🤖 [pic.twitter.com/05ABGxGUqp](https://t.co/05ABGxGUqp?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 9, 2026](https://x.com/testingcatalog/status/2075226270091133348?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Early API partners quoted by Meta framed the release around agentic development and enterprise use. Replit CEO Amjad Masad pointed to the model’s million-token context, multimodal support, search with citations, structured output, parallel tool calling, and an OpenAI-compatible API package. Cline CEO Saoud Rizwan highlighted tool usage and pricing as relevant to scaled coding workloads, while Box VP of AI Products Yashodha Bhavnani said Muse Spark demonstrated enterprise capabilities competitive with frontier models in Box’s internal evaluations. [Source](https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/?ref=testingcatalog.com) ### Mistral unveils Robostral model for robot navigation with RGB camera URL: https://www.testingcatalog.com/mistral-unveils-robostral-model-for-robot-navigation-with-rgb-camera/ Last updated: 2026-07-09T13:46:43.000Z Mistral has introduced Robostral Navigate, an 8-billion-parameter model developed specifically for embodied navigation. This model interprets RGB images and plain-language instructions to autonomously guide robots through complex real-world environments, including offices, residential buildings, commercial spaces, and even outdoor areas. Unlike prior solutions that rely on depth sensors or multiple cameras, Robostral Navigate operates with a single RGB camera and achieves a 76.6% success rate on the R2R-CE benchmark for previously unseen environments. This performance surpasses the best single-camera approach by 9.7 points and outperforms leading multi-sensor systems by 4.5 points. The model runs across a diverse range of robots, including wheeled, legged, and aerial platforms, and is robust to variations in camera hardware. > Announcing Robostral Navigate, our first model for embodied navigation: an 8B robotics navigation model that guides robots to autonomously perform tasks specified with natural language. Single RGB camera. State-of-the-art on R2R-CE. [pic.twitter.com/UlmUsXNxhX](https://t.co/UlmUsXNxhX?ref=testingcatalog.com) > > — Mistral AI (@MistralAI) [July 8, 2026](https://x.com/MistralAI/status/2074856309438980145?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model was trained entirely in simulation, using approximately 400,000 navigation trajectories spanning 6,000 unique scenes. [Mistral](https://www.testingcatalog.com/tag/mistral/) used a specialized data-generation pipeline and a tree-based attention-masking strategy for token-efficient supervised training, significantly reducing training tokens and shortening training time from months to days. The model also benefits from online reinforcement learning, which further increases its performance on navigation tasks. Robostral Navigate is currently available to select partners in manufacturing, logistics, delivery, and hospitality sectors, with plans for broader access as testing continues. Robostral, the company behind this release, has focused on building advanced vision-language models and has designed this system entirely in-house, without relying on existing open-source models, highlighting its commitment to original research and application in robotic navigation. [Source](https://mistral.ai/news/robostral-navigate/?ref=testingcatalog.com) ### OpenAI set to launch ChatGPT Work upgrades today URL: https://www.testingcatalog.com/openai-set-to-launch-chatgpt-work-upgrades-today/ Last updated: 2026-07-09T12:48:21.000Z [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) is expected to unveil ChatGPT Work today, July 9, in a livestream billed as its biggest work-focused update to date. The announcement was teased on LinkedIn and echoed by several OpenAI employees on X, and while public traces remain sparse, early details sketch out what teams and enterprise customers may be getting. > OpenAI announced a new livestream called "ChatGPT Work - Our biggest update for work in ChatGPT" for tomorrow [pic.twitter.com/p9AX0IhMsd](https://t.co/p9AX0IhMsd?ref=testingcatalog.com) > > — Tibor Blaho (@btibor91) [July 8, 2026](https://x.com/btibor91/status/2074965260801654808?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > see you tomorrow [pic.twitter.com/SYvTeDC9Lv](https://t.co/SYvTeDC9Lv?ref=testingcatalog.com) > > — Andrew Ambrosino (@ajambrosino) [July 8, 2026](https://x.com/ajambrosino/status/2074974428723851441?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) One reference points to a pair of settings for resetting a ChatGPT Work environment. The description indicates that all files would be wiped, giving users a clean slate without touching anything stored in the library. That separation suggests a dedicated workspace layer where teams can collaborate with AI on projects, distinct from personal file storage. A plugin labeled "Demo" has also surfaced in the plugin store, most likely intended to showcase what the new solution can do during onboarding or at launch. ![ChatGPT](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/ChatGPT-07-09-2026_02_23_PM-1.jpg) The direction fits OpenAI's recent trajectory of folding Codex ever deeper into ChatGPT. Codex-powered plugins for business roles arrived in June, following [April's introduction of workspace agents](https://www.testingcatalog.com/openai-launched-24-7-always-on-workspace-agents-in-chatgpt/) as an evolution of custom GPTs, so much of what appears in ChatGPT Work is likely to resemble functionality already familiar from Codex: tackling complex projects that require reasoning and sequential execution, with expected gains in context handling, email management, skills, and asset creation. > GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. > > We’re expanding preview access globally now. [pic.twitter.com/Uk5HcfSc2e](https://t.co/Uk5HcfSc2e?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [July 8, 2026](https://x.com/OpenAI/status/2074704958419792299?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Timing matters here too. [GPT-5.6](https://www.testingcatalog.com/openai-launches-gpt-5-6-sol-preview-for-select-partners/) is going public the same day, so it is reasonable to assume ChatGPT Work runs on the new generation of models: Sol, Luna, and Terra. Beyond the work-task demos, there are also mentions of "building workplace [Pets](https://www.testingcatalog.com/openai-prepares-8-interactive-avatars-for-its-codex-app/)". Organization administrators would likely be able to create these, load them with company-specific context, and open them up to everyone in the workspace. Given how attractive ChatGPT Team accounts already are, even for informal groups sharing a subscription, a deeper work layer could considerably widen their appeal. The livestream later today should reveal how much of this lands at launch. ### SpaceXAI launches Grok 4.5 model on Grok Build and APIs URL: https://www.testingcatalog.com/spacexai-launches-grok-4-5-model-on-grok-build-and-apis/ Last updated: 2026-07-09T00:25:56.000Z SpaceXAI has released [Grok 4.5](https://www.testingcatalog.com/spacexai-gearing-up-for-upcoming-grok-4-5-release/), its new flagship model built for coding, agentic workflows, and knowledge work. The model is live now in Grok Build, Cursor on all plans, and the SpaceXAI API console, with EU access not yet active and expected in mid-July. > Announcing Grok 4.5, our first model trained specifically for coding and agents. It was trained with Cursor and offers frontier intelligence at leading speeds and cost efficiency.[https://t.co/i8HpU7w64k](https://t.co/i8HpU7w64k?ref=testingcatalog.com) [pic.twitter.com/oBjGtTsoNc](https://t.co/oBjGtTsoNc?ref=testingcatalog.com) > > — SpaceXAI (@SpaceXAI) [July 8, 2026](https://x.com/SpaceXAI/status/2074915721684086811?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The launch puts Grok 4.5 directly into developer workflows rather than positioning it only as a chatbot upgrade. SpaceXAI says the model was trained alongside Cursor and is now the default model in Grok Build, where it can operate as a coding agent through the CLI, terminal UI, scripts, bots, and Agent Client Protocol integrations. It also runs inside Office add-ins for Word, PowerPoint, and Excel, where the company says it can create spreadsheet models, slide diagrams, and document drafts. > We've partnered with SpaceXAI to train Grok 4.5. > > It’s our most powerful model yet and the first we've built for more than software engineering. [pic.twitter.com/U4B8Tedl34](https://t.co/U4B8Tedl34?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [July 8, 2026](https://x.com/cursor%5Fai/status/2074915744999969059?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Grok 4.5 supports text and image input, text output, a 500,000-token context window, function calling, structured outputs, web search, X search, code execution, and configurable reasoning. Developers can set reasoning effort to low, medium, or high, with high used by default and reasoning not disableable. API pricing is listed at $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens, with higher-context pricing applying above 200,000 tokens. SpaceXAI is pitching the model around real engineering tasks. In its published benchmarks, Grok 4.5 scored: 1. 62.0% on DeepSWE 1.0 2. 53% on DeepSWE 1.1 3. 83.3% on Terminal Bench 2.1 4. 64.7% on SWE Bench Pro The company also claims the model serves at 80 tokens per second and uses 15,954 output tokens on average for SWE Bench Pro tasks, which it frames as about 4.2 times fewer tokens than Opus 4.8 max in the same comparison. > SpaceXAI just released Grok 4.5, and it ranks #4 on GDPval-AA v2 with an Elo of 1543 - behind only the latest Claude releases from Anthropic on real-world agentic knowledge work tasks > > Grok 4.5 achieved this score at a cost of $0.49 per GDPval task to sit clearly on the Pareto… [pic.twitter.com/XzK3xq83hf](https://t.co/XzK3xq83hf?ref=testingcatalog.com) > > — Artificial Analysis (@ArtificialAnlys) [July 8, 2026](https://x.com/ArtificialAnlys/status/2074942097158021371?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model was trained across tens of thousands of NVIDIA GB300 GPUs, with data filtering, deduplication, quality scoring, domain-focused selection, and reinforcement learning over hundreds of thousands of tasks. [SpaceXAI](https://www.testingcatalog.com/tag/grok/) says the RL process targeted multi-step software engineering and technical work, with automated and model-based grading plus long-running agentic rollouts. [Source](https://x.ai/news/grok-4-5?ref=testingcatalog.com) ### OpenAI rolls out GPT-Live voice for ChatGPT on web and mobile URL: https://www.testingcatalog.com/openai-rolls-out-gpt-live-voice-for-chatgpt-on-web-and-mobile/ Last updated: 2026-07-09T00:11:24.000Z OpenAI has launched GPT-Live, a new family of voice models for ChatGPT Voice, with GPT-Live-1 and GPT-Live-1 mini rolling out globally starting July 8, 2026\. The update moves ChatGPT Voice beyond turn-based conversations by using a [full-duplex architecture](https://www.testingcatalog.com/openai-prepares-bidirectional-voice-mode-for-rollout-on-chatgpt/), allowing the system to listen and speak simultaneously, handle interruptions, pause when the user is thinking, and continue a conversation while deeper tasks run in the background. > Introducing GPT-Live, a new generation of voice models for natural human-AI interaction. > > Rolling out in ChatGPT starting today. > > You’ll want to turn the sound on for this one. [pic.twitter.com/WzoQFvA5ir](https://t.co/WzoQFvA5ir?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [July 8, 2026](https://x.com/OpenAI/status/2074907025537224840?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The launch targets everyday ChatGPT users first. GPT-Live-1 becomes the default voice model for Go, Plus, and Pro users, while GPT-Live-1 mini becomes the default for Free users. The rollout covers ChatGPT on iOS, Android, and web, while Business, Enterprise, and Edu workspaces are not included at launch. OpenAI says API access is planned soon, with developers and enterprises able to register interest. GPT-Live can delegate harder requests to frontier models behind the scenes. At launch, that means GPT-5.5 handles search, reasoning, or more complex work while GPT-Live keeps the spoken conversation active. Users can also select Instant, Medium, or High intelligence levels where available, trading response speed for deeper reasoning. ChatGPT Voice now supports visual cards for topics such as weather, stocks, and sports, as well as web search, memory, text, and images in supported accounts. > A new GPT Live 1 voice mode is rolling out super fast. It also comes with an effort selector and updated quick setting menu on mobile. > > Testing time 👀 [https://t.co/uB4nFGV5K8](https://t.co/uB4nFGV5K8?ref=testingcatalog.com) [pic.twitter.com/7ZRuQ4BSu6](https://t.co/7ZRuQ4BSu6?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 8, 2026](https://x.com/testingcatalog/status/2074955436810240376?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) This is a direct replacement path for OpenAI’s earlier voice systems. The original Standard Voice Mode chained speech-to-text, an LLM, and text-to-speech, which added latency. Advanced Voice Mode processed audio using a single model but still relied on discrete turns. GPT-Live changes that model by continuously processing input and output, allowing it to decide many times per second whether to speak, listen, pause, interrupt, or call a tool. The initial release has clear limits. Live does not support video or screen sharing at launch, and it is not initially available in Temporary Chats, the ChatGPT desktop app, Work, Codex, or custom GPTs. Advanced Voice Mode remains available for eligible users who need mobile video or screen sharing. [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) is positioning GPT-Live as the foundation for longer, more agentic voice work, not just casual chat. The company says more than 150 million people use ChatGPT Voice and Dictation each week, giving this rollout immediate scale across consumer ChatGPT. [Source](https://openai.com/index/introducing-gpt-live/?ref=testingcatalog.com) ### ByteDance debuts Seedream 5.0 Pro with advanced reasoning URL: https://www.testingcatalog.com/bytedance-debuts-seedream-5-0-pro-with-advanced-reasoning/ Last updated: 2026-07-09T00:05:16.000Z ByteDance Seed has launched Seedream 5.0 Pro, a multimodal image-creation model designed for production design work rather than one-shot image output. The model targets creators, designers, marketers, educators, product teams, and developers who need dense visual layouts, localized text, controlled edits, and reusable design assets from prompt and reference inputs. The new Pro model brings four core upgrades: complex information visualization, precision editing, realistic imagery and portrait textures, and native multilingual input and generation. ByteDance says it can turn data, concepts, and long text into professional layouts, including infographics, posters, UI mockups, educational visuals, and structured commercial assets. Compared with earlier Seedream models, the company points to stronger image-text alignment, structural coherence, text rendering, and visual aesthetics. > Seedream 5.0 Pro is UNLIMITED on Magnific in 1.5K resolution > > Precise image generation powered by [@byteplusglobal](https://x.com/BytePlusGlobal?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > > → Generates text natively in 14 languages > → Full infographics in one go > → Precision editing straight into your design workflow > > Now available on Magnific [pic.twitter.com/VnuIFxF8pc](https://t.co/VnuIFxF8pc?ref=testingcatalog.com) > > — Magnific (@magnific) [July 8, 2026](https://x.com/magnific/status/2074843521853636609?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Seedream 5.0 Pro can use point selection, lasso selection, box selection, sketches, color swatches, material references, and multi-image inputs to edit a specific region rather than regenerating the entire image. It also supports layer separation, letting a poster be split into editable assets such as text, subject, background, and decorations, with occluded background areas restored during the process. > Seedream 5.0 Pro is now available on fal > > Region-precise editing that changes one element and leaves the rest untouched > Advanced prompt understanding and native text in 14 languages > Suitable for structured designs, posters, product mockups, UI-style layouts, charts [pic.twitter.com/VKJC1UdrdS](https://t.co/VKJC1UdrdS?ref=testingcatalog.com) > > — fal (@fal) [July 8, 2026](https://x.com/fal/status/2074846830198722944?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For realistic outputs, ByteDance is pushing the model toward physical lighting, material behavior, skin texture, photographic motion, and multi-person compositing. The company shows examples covering glass reflections, architectural photography, cinematic portraits, panning shots, and group photos built from multiple source images. For global use cases, Seedream 5.0 Pro supports more than 10 languages, including Chinese, English, French, German, Russian, Japanese, Korean, Spanish, and Arabic, with support for right-to-left layouts and accents. > A new seedream-5-0-pro image generation model from ByteDance is now available on BytePlus APIs for testing. > > \> It supports text-to-image, single image-to-image, and multi-reference image-to-image. [https://t.co/PdEIQAgSfW](https://t.co/PdEIQAgSfW?ref=testingcatalog.com) [pic.twitter.com/IPsMvACMA2](https://t.co/IPsMvACMA2?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 8, 2026](https://x.com/testingcatalog/status/2074871631407628360?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Developer access is also taking shape through BytePlus ModelArk. BytePlus documentation lists seedream-5-0-pro as a new image generation model that can generate a single image or produce 2 to 10 reference images from a text prompt. The same documentation section was updated on July 8, 2026, indicating an API-side rollout around the launch window. ByteDance Seed is the company’s foundation model research unit, working across LLMs, infrastructure, vision, speech, multimodal systems, AI for science, robotics, and responsible AI. Seedream 5.0 Pro sits inside its GenMedia push alongside models such as Seedance 2.0 and Seedream 5.0 Lite, moving ByteDance deeper into creator tools where image generation, editing, layout, and multilingual production are converging. [Source](https://seed.bytedance.com/en/seedream5%5F0%5Fpro?ref=testingcatalog.com) ### Meta launches Muse Image across its apps and previews Muse Video URL: https://www.testingcatalog.com/meta-launches-muse-image-across-its-apps-and-previews-muse-video/ Last updated: 2026-07-09T00:01:48.000Z Meta has started rolling out Muse Image, its first image-generation model from Meta Superintelligence Labs, inside Meta AI, turning the assistant into a visual creation tool across the company’s consumer apps. The model is now available in the Meta AI app, with Instagram Stories support in the US and WhatsApp image generation in select countries. Facebook, Messenger, more Instagram and WhatsApp surfaces, and advertiser access through Advantage+ creative are next in line. > META 🔥: Muse Image, the first image-gen model from MSL, is now available on Meta AI. > > \> It uses advanced reasoning to understand complex prompts, seamlessly blending multiple photos into high-quality creations you can download and share anywhere. > > Users can also access Presets,… [https://t.co/GgV5F09zta](https://t.co/GgV5F09zta?ref=testingcatalog.com) [pic.twitter.com/DX1Zj5DFjP](https://t.co/DX1Zj5DFjP?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 7, 2026](https://x.com/testingcatalog/status/2074571946642059555?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Muse Image is built for prompt-based image creation, photo editing, multi-reference composition, room redesigns, personalized presets, and social creation using Instagram context. Users can start with text, existing photos, suggested prompts, @-mentioned public Instagram accounts, or sketches drawn directly on top of an image. Meta says the model can render cleaner text within visuals, generate QR codes and plots via coding tools, and use web search to ground images in current or factual context. The release makes Meta’s image strategy more tightly tied to its own social graph than standalone image generators. Muse Image can draw from public Instagram photos when a user tags an account, while Instagram users have the option to turn off this type of AI reuse. That puts the launch directly inside creators’ existing workflows, from story effects and chat images to event graphics, room makeovers, product concepts, and small-business marketing assets. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Meta-AI-07-09-2026_01_39_AM--1-.jpg) Under the hood, Meta frames Muse Image as an agentic media model rather than a basic text-to-image system. It can plan before generation, use search and coding tools, self-refine outputs, scale reasoning at inference time, and work with Muse Spark, Meta’s earlier reasoning model. Meta also says Muse Image ranks No. 2 on Arena for text-to-image, single-image editing, and multi-image editing as of July 5, 2026, behind GPT Image 2 in the text-to-image leaderboard and ahead of models from Reve, Google, Microsoft AI, xAI, and others. ![Meta](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Meta-AI-07-09-2026_01_39_AM.jpg) Meta is also adding Content Seal, an invisible watermark for images generated in Meta AI. The company says the signal is designed to remain detectable after cropping, compression, resizing, or screenshots, and it is previewing a detection tool for checking whether an image carries the watermark. Meta says the same provenance system is planned for video later. > META 🔥: In addition to Muse Image, Meta has announced the Muse Video model! > > \> Muse Video is built upon the same pretraining base as Muse Image to deliver exceptional visual fidelity with native audio support. > > \> Muse Video is coming soon to creators and Meta AI. > > Loads of good… [https://t.co/E60FzrvriG](https://t.co/E60FzrvriG?ref=testingcatalog.com) [pic.twitter.com/EpYFQRJ3x2](https://t.co/EpYFQRJ3x2?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 7, 2026](https://x.com/testingcatalog/status/2074594946955297165?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [Meta](https://www.testingcatalog.com/tag/meta/) is positioning Muse Image as the second major step in its Muse roadmap after Muse Spark, which began powering Meta AI earlier this year across the [Meta AI](https://meta.ai/?ref=testingcatalog.com) app, WhatsApp, Instagram, Facebook, Messenger, Threads, and AI glasses. With Muse Video already previewed and coming later to creators and Meta AI, Meta is now moving from assistant answers into social media production, ad creative, and generated visuals inside the apps where its users already post and message. [Source](https://about.fb.com/news/2026/07/introducing-muse-image-meta-ai/?ref=testingcatalog.com) ### Anthropic brings Claude Cowork to web and mobile for Max users URL: https://www.testingcatalog.com/anthropic-brings-claude-cowork-to-web-and-mobile-for-max-users/ Last updated: 2026-07-08T14:00:21.000Z Anthropic is expanding Claude Cowork to web and [mobile platforms](https://www.testingcatalog.com/anthropic-prepares-cowork-support-for-mobile-apps/), transforming its agentic work tool from a desktop-first product into a cross-device workspace for extended knowledge work. The rollout begins with Max users and will gradually extend to more plans over the coming weeks, while web and mobile access remains in beta. Claude Cowork enables users to delegate tasks to Claude across connected files, calendar, email, messaging apps, the web, and other tools. The primary enhancement is continuity: a task can initiate on a desktop, continue in the background after the laptop is closed, and be reviewed or redirected from a phone. Scheduled tasks can now operate without a device being online, advancing Cowork towards an always-available work agent rather than a chat session tied to an active screen. > Claude Cowork is coming to mobile and web. > > Hand Claude a task at your desk and pick up the finished work from your phone. Close the laptop and Claude keeps going. > > Beta is rolling out over the next several weeks starting with the Max plan, with more plans to follow. [pic.twitter.com/W4WrgN9TrG](https://t.co/W4WrgN9TrG?ref=testingcatalog.com) > > — Claude (@claudeai) [July 7, 2026](https://x.com/claudeai/status/2074525815820169320?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The feature is accessible through the Claude home screen on the web and from the sidebar in the Claude apps for iOS and Android. The desktop remains the full version, as it can use local files, local connectors, browser access, and computer access through the installed app. Web and mobile users can start, steer, resume, and review tasks; use connectors, skills, plugins, and scheduled tasks; manage projects; and preview files Claude creates. Live artifacts remain desktop-only for now. Anthropic is positioning Cowork for everyday business operations and content workflows, not just coding. The company reports that more than 90% of Cowork usage is outside software development, with significant use cases including: 1. Operations 2. Finance 3. Legal 4. Sales 5. Marketing 6. Research synthesis 7. Deck creation 8. Contract review 9. Variance reporting A separate usage analysis sampled 1.2 million anonymized Cowork sessions from late May and identified business process and operations as the largest category at 33.4%, followed by content creation and copywriting at 16.4%. Human approval remains integral to the product design. Claude can continue working in the background, but when it reaches a decision that requires user judgment, it requests input and can send the question to the phone. Anthropic emphasizes that user review is still necessary before outputs are sent or shipped, maintaining Cowork as a delegated-work system rather than a fully autonomous publishing or execution layer. Anthropic is also extending doubled five-hour Cowork usage limits through August 5 for eligible Pro, Max, Team, and legacy seat-based Enterprise users. Free users and consumption-based Enterprise seats are not included. The promotion applies only to Cowork, not regular Claude chat or Claude Code. [Source](https://claude.com/blog/cowork-web-mobile/?ref=testingcatalog.com) ### SpaceXAI launches 21 new Grok Voices on xAI Console URL: https://www.testingcatalog.com/spacexai-launches-21-new-grok-voices-on-xai-console/ Last updated: 2026-07-07T12:16:00.000Z xAI is rolling out 21 new flagship voices for Grok Voice, expanding its built-in voice roster from five to 26 and enhancing its capabilities in multilingual voice agents, text-to-speech, and no-code agent creation. These new voices are now available across the real-time Voice Agent API, Text-to-Speech API, and Grok Voice Agent Builder, targeting developers, operators, support teams, sales teams, educators, advertisers, and businesses building production voice agents. > Grok Voice now has 25+ flagship voices that each natively support 25+ languages! [https://t.co/0QDsDEmsO3](https://t.co/0QDsDEmsO3?ref=testingcatalog.com) [pic.twitter.com/fNcmCd7bvk](https://t.co/fNcmCd7bvk?ref=testingcatalog.com) > > — Eden Chan (@edenchan) [July 7, 2026](https://x.com/edenchan/status/2074302886604046356?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The release introduces voices such as Carina, Zagan, Helix, Orion, Luna, Iris, Altair, Zenith, Perseus, Helios, Lux, Kepler, Rigel, Cosmo, Celeste, Ursa, Sirius, Lumen, Castor, Naksh, and Atlas. xAI positions them around specific use cases, including: 1. Support 2. Characters 3. Commentary 4. Advertising 5. Education 6. Wellness 7. Narration 8. Sales 9. Assistants 10. Podcasts 11. Audiobooks The original five voices, Ara, Eve, Leo, Rex, and Sal, have also been retrained for more natural pacing, phrasing, and emphasis. The voice system operates across Grok’s real-time and TTS stack. Developers can stream audio and text over WebSocket for voice assistants, phone agents, and voice systems. Session settings support built-in or custom voices, server-side voice activity detection, tools, web search, X search, file search, MCP, function calls, pronunciation replacements, and adjustable speech speed. The current flagship real-time model is Grok Voice Think Fast 1.0, while the legacy Grok Voice Fast 1.0 model is marked as deprecated. For text-to-speech, xAI supports expressive speech tags such as pauses, laughter, breathing, whispering, volume, pitch, and speed changes, singing, and emphasis. The TTS API accepts up to 15,000 characters, supports MP3 by default at 24 kHz and 128 kbps, and can return character-level timing metadata. Official documentation lists 20 supported TTS languages, while xAI’s announcement states the new voices are natively multilingual across Grok Voice’s 25+ languages. The launch is closely tied to xAI’s new [Voice Agent Builder](https://www.testingcatalog.com/icymi-xai-debuts-grok-voice-agent-builder-for-enterprises/), a beta no-code tool that lets users create voice agents in about two minutes. The builder includes features such as telephony, knowledge retrieval, tools, guardrails, MCP support, observability, SIP number support, call recordings, transcripts, and tool-use logs. xAI states that agents are billed at the API rate of $0.05 per minute of audio, with voice included and no separate platform fee, while telephony on a free-provisioned number costs an additional $0.01 per minute. The company behind the release is xAI, which focuses on frontier reasoning, real-time voice, and generative media. The voice rollout expands [Grok](https://www.testingcatalog.com/tag/grok/) from a chatbot and API model family into a broader production stack for call centers, voice assistants, and brand-controlled speech experiences. Custom Voices are also significant: xAI allows users to clone a voice from a reference clip up to 120 seconds long and use it across TTS and real-time voice APIs. However, Custom Voices are currently available only in the United States except Illinois, with API creation gated to Enterprise teams. [Source](https://x.ai/news/new-flagship-voices?ref=testingcatalog.com) ### SpaceXAI gearing up for upcoming Grok 4.5 release URL: https://www.testingcatalog.com/spacexai-gearing-up-for-upcoming-grok-4-5-release/ Last updated: 2026-07-07T07:45:58.000Z References to Grok 4.5 have begun surfacing across the products now grouped under SpaceXAI, the banner Elon Musk adopted after folding xAI into SpaceX earlier this year. Two early signals point toward a public launch that may not be far off. > We are now [@SpaceXAI](https://x.com/SpaceXAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com). [pic.twitter.com/ema66xDWC9](https://t.co/ema66xDWC9?ref=testingcatalog.com) > > — SpaceXAI (@SpaceXAI) [July 6, 2026](https://x.com/SpaceXAI/status/2074214064746832060?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The first is a **feature flag in the Grok web client** that swaps alternate copy into the account upgrade page, where Grok 4.5 looks set to be named; it sits dormant for now and promises nothing about timing, but the scaffolding is in place. The second is the **model's appearance in Grok Build CLI**, the company's terminal coding agent, within the same flow that prompts users to move up to the Heavy plan before trying it. > 🚨BREAKING🚨 Its happening!. Grok Build is showing that Grok 4.5 is Here and prompting users to upgrade to SuperGrok Heavy to access it! > > You may want to do that quickly to acces the $99/mo promo that is still active right now! [https://t.co/FHdmi4TOKH](https://t.co/FHdmi4TOKH?ref=testingcatalog.com) [pic.twitter.com/WepllBkEzI](https://t.co/WepllBkEzI?ref=testingcatalog.com) > > — ️️️️ ️ᅠ‏️️️️ ️ᅠ️️️️ ️️️️️ ️ᅠ (@blankspeaker) [July 6, 2026](https://x.com/blankspeaker/status/2074195091653378476?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Grok 4.5 is expected to run on a fresh V9 foundation with roughly 1.5 trillion parameters, making it the largest model the company has shipped. Musk has described its performance as close to Opus-level without specifying which version of Anthropic's flagship he meant. That framing carries weight, because the claim rests on internal evaluations at SpaceX and Tesla, where the model has reportedly been running for a couple of weeks, rather than any public benchmark. Independent numbers do not yet exist. > Grok 4.5, based on our 1.5T V9 foundation model, with Cursor data added in supplemental training, is now in private beta at SpaceX & Tesla. Early evals show performance close to, perhaps exceeding Opus. > > RL is continuing to significantly improve the model, and the Grok Build… > > — Elon Musk (@elonmusk) [June 28, 2026](https://x.com/elonmusk/status/2071184354756477041?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Access looks set to follow the pattern of earlier releases and of grok-build itself: 1. Heavy subscribers first 2. Then a wider opening to SuperGrok tiers For [SpaceXAI](https://www.testingcatalog.com/tag/grok/), the model is a test of whether the compute and capital gained through the merger translate into a system that narrows the gap with Anthropic and OpenAI in reasoning and coding, rather than on parameter count alone. The dormant upgrade copy suggests the commercial pieces are being readied ahead of any performance reveal. Whether the release lands on time, against a target that already slipped past late spring, is the open question. ### ByteDance set to launch Seedance 2.5 with 3-minute AI video output URL: https://www.testingcatalog.com/bytedance-set-to-launch-seedance-2-5-with-3-minute-ai-video-output/ Last updated: 2026-07-04T19:35:36.000Z ByteDance appears close to releasing Dreamina Seedance 2.5, with current rumors pointing to a **possible July 9 launch**. Public Dreamina and CapCut pages already reference the model, while third-party reports suggest an early-July rollout window rather than a confirmed date. Availability is expected across Dreamina, CapCut, and partner platforms that already work with ByteDance’s video stack. > Coming soon: Dreamina Seedance 2.5 is arriving on CapCut. > > Seamless generation and editing. Up to 50 multimodal references. 30-second scenes in one shot. Finer creative control. More reliable results. It's built to make creating faster, smoother, and more intuitive. > > Whether… [pic.twitter.com/hkJnP3zg6Y](https://t.co/hkJnP3zg6Y?ref=testingcatalog.com) > > — CapCut (@capcutapp) [July 4, 2026](https://x.com/capcutapp/status/2073261464065122562?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The main new feature will be a move from short clips to 30-second standard video generation. Dreamina pages describe a long video workflow that can generate 30-second scenes, 90-second drafts, and 180-second outputs, making three-minute AI video the biggest claimed shift for this release. The open question is not only about duration but also about whether the model can keep character identity, motion, camera logic, and prompt intent stable across extensions. > You can create [cinematic videos](https://dreamina.capcut.com/create/cinematic-video?ref=testingcatalog.com) up to 30 seconds in standard mode or extend them to 180 seconds with the beta long-video mode. ![Dreamina website](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Official-Seedance-2-5-AI-Video-Generator-30s-Clips-50-References-07-04-2026_09_28_PM.jpg) [Dreamina website](https://dreamina.capcut.com/seedance/seedance-2-5?ref=testingcatalog.com) For creators, advertisers, anime editors, social video accounts, and AI filmmakers, this would turn Seedance from a short-shot generator into a tool for longer sequences. The likely location is within Dreamina’s Seedance workflow, CapCut’s creator tools, and partner apps that use ByteDance’s model access. The evidence appears to come from public-facing product pages and platform copy, while the July 9 timing remains rumor-level until ByteDance posts a dated release notice. [Join Dev Mode ](https://discord.com/invite/devmode?ref=testingcatalog.com) The company behind the model is ByteDance, which has been tying AI video to its creator ecosystem through Dreamina, CapCut, TikTok-adjacent production flows, and API distribution. Seedance 2.0 is still positioned around motion stability, multimodal references, and audio-video generation, with listed durations of 4 to 15 seconds on BytePlus. Seedance 2.5 would push that strategy toward longer commercial storytelling and creator workflows, while Google’s Veo and [Gemini Omni](https://www.testingcatalog.com/google-rolls-out-gemini-omni-ai-for-video-generation-and-editing/) remain the clearest competitive pressure after OpenAI discontinued Sora’s web and app product on April 26, 2026. [Source](https://x.com/MarsForTech?ref=testingcatalog.com) ### Meta prepares Scheduled tasks for Meta AI users on web URL: https://www.testingcatalog.com/meta-prepares-scheduled-tasks-for-meta-ai-users-on-web/ Last updated: 2026-07-04T17:59:19.000Z [Meta](https://www.testingcatalog.com/tag/meta/) appears to be building a scheduled tasks section for Meta AI on the web, where users could set up and manage recurring instructions instead of issuing every prompt by hand. Spotted in recent builds and not yet live for anyone, this feature would allow users to create jobs that run on a set cadence, such as a morning news digest, a weekly summary, or a periodic check on a topic, with the results delivered back inside the assistant. The naming leans toward the industry-standard "scheduled tasks," and the groundwork echoes an earlier [Tasks](https://www.testingcatalog.com/meta-ai-redies-avacado-manus-agent-and-openclaw-integration/) effort seen some time ago, suggesting the next turn of that idea rather than a fresh start. What remains unclear is whether these jobs can trigger on events, reach connected apps, or push notifications when they finish. The first beneficiaries would be web users who want Meta AI to act on its own, and Meta itself, which has spent much of this year closing the distance with rivals whose assistants already run scheduled tasks, ChatGPT among them. Recurring, unattended work moves the assistant from a question box toward something closer to an agent, the direction every major lab is pursuing. > First, Mark was clearly talking about the industry’s progress on agentic capabilities on the whole. > > But, while we’re on the topic: Our next Muse Spark update is coming soon. Big improvements in coding and agentic capabilities to be more competitive with other leading models.… [https://t.co/uTjx8sZM2A](https://t.co/uTjx8sZM2A?ref=testingcatalog.com) > > — Alexandr Wang (@alexandr\_wang) [July 3, 2026](https://x.com/alexandr%5Fwang/status/2072848108342677597?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The timing sits beside louder model news. Alexandr Wang, Meta's chief AI officer, told staff at a July 2 town hall that the company's next model, codenamed Watermelon and still in training, has caught up with OpenAI's GPT-5.5 on internal benchmarks and draws far more compute than Muse Spark, the April model known internally as Avocado. > ya, really big ones actually > > — Alexandr Wang (@alexandr\_wang) [July 3, 2026](https://x.com/alexandr%5Fwang/status/2072852259613151511?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Asked about matching Anthropic's Opus on coding, he pointed not to Watermelon but to a forthcoming Muse Spark update, due "pretty soon." No release date or public benchmarks exist yet, and Zuckerberg struck a more guarded note, citing a three- to six-month window for returns. How fast Watermelon ships will decide whether it meets today's frontier or arrives to face the next one. ### Google tests new Gemini Inbox section for Workspace triage URL: https://www.testingcatalog.com/google-tests-new-gemini-inbox-section-for-workspace-triage/ Last updated: 2026-07-04T15:48:41.000Z Google appears to be building a dedicated inbox section inside the Gemini app for Business and Workspace customers. This feature has been spotted in recent builds but is not yet available to anyone. The layout centers on three filters that allow a person to sort items to follow up on, review what has been marked done, and check work that is ready for review. The framing points at an Inbox Zero workflow that pulls messages out of Gmail and into the Gemini surface itself, rather than layering help on top of the mailbox. The more telling signal is the "Needs review" filter, which suits a proactive agent that triages email and data from connected sources, then files what it finds into a structured to-do list for a person to walk through. This pattern already runs through some other Google apps: 1. Daily Brief compiles urgent mail and calendar items into a morning summary. 2. [Gemini Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/) operates as a background agent that can archive newsletters and surface follow-ups. 3. The Gmail-based AI Inbox, shown to testers earlier this year, lifts deadlines and tasks to the top. Placing this view inside [Gemini](https://www.testingcatalog.com/tag/gemini/) aligns with the company's wider direction. Google has spent the past year transforming Gemini from a chatbot into a worker, adding a macOS app, browser-driven agent runs, and the no-code automation builder, now called Workspace Studio, which is accessible from Gmail. Folding email triage, task tracking, and those automations into one panel would push Gemini toward a consolidated desktop workspace, a single place where Computer Use and Browser control sit beside the mailbox. Whether that becomes a true super app for Workspace remains open, though rival agent platforms are moving in the same direction. ### OpenAI might be preparing GPT-5.6 for next week's release URL: https://www.testingcatalog.com/openai-might-be-preparing-gpt-5-6-for-next-weeks-release/ Last updated: 2026-07-04T14:08:03.000Z OpenAI has moved **GPT-5.6** into a narrow preview, and fresh signals in the Codex app hint at how the company plans to surface it once the gate lifts. The family, unveiled on June 26, splits into three tiers: Sol as the flagship, Terra as a mid-cost option, and Luna as the fastest and cheapest. For now, the trio reaches only vetted partners through Codex and the API, with no place in ChatGPT. What stands out in recent Codex builds is a **reworked reasoning-effort control**, rendered as a slider rather than preset buttons. This tracks with a confirmed part of GPT-5.6: a new "max" setting that gives Sol more room to work through long problems, plus a subagent-driven "ultra" mode for heavier jobs. A slider would hand developers one control to trade speed against depth. The layout resembles the reasoning selector already on Anthropic's Claude Code desktop client, though it could shift before anything ships. References to real-time voice, present in earlier builds, also appear to have been removed from the current app, with no clarity on whether the capability will still be developed for Codex. The developers and enterprises leaning on Codex for agentic coding stand to feel this first, and the timing is pointed. Anthropic's Fable 5, restored globally on July 1, **will no longer be bundled into subscription plans on July 7** and will move to usage-based credits, prompting some users to weigh alternatives just as OpenAI courts them. The Sol, Terra, and Luna labels mark a shift toward durable capability tiers that each advance on their own schedule, a rethink of how the company names and prices its lineup. Broad access is **rumored for the same window** but stays tied to voluntary US government review under a recent cybersecurity order, so approvals, not a fixed calendar, will decide when GPT-5.6 opens to everyone. ### Mistral releases Leanstral 1.5 open model for proof engineering URL: https://www.testingcatalog.com/mistral-releases-leanstral-1-5-open-model-for-proof-engineering/ Last updated: 2026-07-04T12:56:43.000Z Mistral AI has released Leanstral 1.5, a new open-source code agent model built for Lean 4 formal proof engineering, automated theorem proving, and autoformalization. The model is available as `labs-leanstral-1-5` through Mistral’s Labs API, in Mistral Vibe, and as downloadable weights on [Hugging Face](https://huggingface.co/mistralai/Leanstral-1.5-119B-A6B?ref=testingcatalog.com) under an Apache-2.0 license. The release targets researchers, proof engineers, developers working with formal methods, and teams exploring verified software. Mistral’s documentation lists Leanstral 1.5 with 119B total parameters, 6.5B active parameters, a 256k context window, text and image input, text output, and $0 pricing in Labs. The Labs listing also notes that the model is scheduled for retirement on September 30, 2026, indicating a limited-window experimental deployment rather than a permanent production endpoint. > Despite being primarily trained on math, Leanstral demonstrates impressive code verification capabilities, discovering previously unknown bugs in open-source repositories. > > We built an automated pipeline where Aeneas translates Rust code to Lean and Leanstral infers the user… > > — Mert Ünsal (@mertunsal2020) [July 3, 2026](https://x.com/mertunsal2020/status/2073046292734111846?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Mistral states that Leanstral 1.5 raises the ceiling for machine-checked reasoning in Lean 4\. The company reports that the model fully saturates miniF2F, solves 587 of 672 PutnamBench problems, reaches 87% on FATE-H and 34% on FATE-X, and lifts FLTEval pass@8 from 31.9 to 43.2\. Mistral also mentions that the model can continue working on very long proof attempts, citing a case where Leanstral processed more than 2.7 million tokens across 22 context compactions while proving AVL-tree time-complexity guarantees. The system was trained through mid-training, supervised fine-tuning, and reinforcement learning with CISPO. In one training environment, Leanstral receives theorem statements, submits proofs, reads Lean compiler feedback, and revises until the proof compiles or the attempt budget ends. In another, it acts like a developer inside a raw filesystem, editing files, running bash commands, using the Lean language server, building helper lemmas, and completing partial proofs inside real repositories. Mistral is also positioning the model beyond academic math. In a Rust verification pipeline using Aeneas and Lean, the company says Leanstral generated correctness properties, tried to prove them, then attempted to prove their negation when proofs failed. Across 57 repositories, Mistral claims that this process flagged 47 violations, 11 genuine bugs, and 5 bugs that had not previously been reported on GitHub. ![Mistral](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Leanstral-charts_Z1KXfze.webp) This move builds on the first Leanstral release from March 2026, which Mistral described as its first open-source code agent for Lean 4 proof engineering. The older `labs-leanstral-2603` model is listed as retired on June 30, 2026, and replaced by Leanstral 1.5\. Early developer response is already forming around the tooling layer, with the maintainer of OpenATP stating that the package would be updated to point at Leanstral 1.5. For [Mistral](https://www.testingcatalog.com/tag/mistral/), the release aligns with its broader push to offer full-stack AI systems spanning frontier models, developer tools, applications, and compute. The company frames Leanstral as part of its open model strategy and as a technical bet on AI agents that not only generate code but can also prove properties about code and mathematics through formal verification. [Source](https://mistral.ai/news/leanstral-1-5/?ref=testingcatalog.com) ### ICYMI: xAI debuts Grok Voice Agent Builder for Enterprises URL: https://www.testingcatalog.com/icymi-xai-debuts-grok-voice-agent-builder-for-enterprises/ Last updated: 2026-07-04T12:41:51.000Z xAI has transitioned Grok Voice from a model/API play into a comprehensive no-code voice agent platform, introducing the Voice Agent Builder in beta. This platform is designed for operators and developers who aim to create production phone agents without having to manually assemble components such as speech-to-text, reasoning, text-to-speech, telephony, tools, guardrails, MCP support, and observability. Available through the xAI Console, it promises to create a personalized voice agent in under two minutes using a plain-language call flow, uploaded documents, connected tools, and a browser-based test call. > Introducing Voice Agent Builder: a no-code platform to create human-like voice agents with Grok Voice. > > Available today at $0.05 / min.[https://t.co/kUkF7zqvfR](https://t.co/kUkF7zqvfR?ref=testingcatalog.com) [pic.twitter.com/OCIq1oDYar](https://t.co/OCIq1oDYar?ref=testingcatalog.com) > > — xAI (@xai) [July 1, 2026](https://x.com/xai/status/2072342803787702422?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The platform is targeted at high-volume call workflows such as customer support, sales, lead qualification, reception, and scheduling. Agents can access knowledge bases, search documents during calls, and connect to services like Gmail, Google Calendar, Outlook, Linear, Notion, OneDrive, Google Drive, APIs, X search, web search, and remote MCP servers. They can also transfer callers to a human when necessary. Each call can be recorded, transcribed, replayed, and inspected, with tool usage visible for review. xAI is positioning this launch against the typical voice-agent stack that combines separate speech recognition, a language model, and speech synthesis. The company claims that Grok Voice employs a more integrated speech-to-speech path, offering sub-second latency, support for over 25 languages, and handling of noisy phone audio, accents, interruptions, and callers who change direction mid-call. xAI also asserts that Grok Voice Think Fast 1.0 leads its τ-voice Bench table with a score of 67.3%, surpassing Gemini 3.1 Flash Live at 43.8% and GPT Realtime 1.5 at 35.3% on the same benchmark. Voice configuration includes over 80 built-in voices in the Builder experience, as well as brand voice cloning from approximately 2 minutes of audio. Businesses have the option to use a free xAI-provisioned phone number, bring an existing number through SIP, or connect their own client over WebSocket. Pricing is structured around xAI’s real-time voice rate of $0.05 per minute of audio, with an additional $0.01 per minute for telephony on an xAI-provisioned number. xAI states that voices are included and there is no separate platform fee. [Source](https://x.ai/news/grok-voice-agent-builder?ref=testingcatalog.com) ### Condense launches proxy to cut AI coding agent bills by up to 72% URL: https://www.testingcatalog.com/condense-launches-proxy-to-cut-ai-coding-agent-bills-by-up-to-72/ Last updated: 2026-07-13T15:12:03.000Z Condense.chat has opened public access to a context-compression proxy for coding agents, a system that sits between an agent and the upstream model, shrinking each request before billing. It targets a line item most teams never inspect. In a working agent loop, the model is re-sent the system prompt plus the whole conversation on every turn, so by the middle of a long session, it has re-read the same early context hundreds of times. Across twelve real coding sessions and 18,333 assistant turns, Condense puts cache reads at 67.7 percent of a typical bill, the cost the proxy is built to remove. > We're giving you 100M free tokens to prove your agent is wasting context. > > Most of what your agent sends upstream is dead weight. We built state-of-the-art compaction to strip it, and today we published the receipts. > > One real session, run to full depth, benchmarked against… [pic.twitter.com/40ZdASrHoj](https://t.co/40ZdASrHoj?ref=testingcatalog.com) > > — condense.chat (@densechat) [July 3, 2026](https://x.com/densechat/status/2073029082283974737?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The proxy runs two compression models in sequence. Helene 1, an extractive model, scores every token and keeps the survivors verbatim as a strict subset of what the agent saw, stripping content before it reaches cache. Adeline 1 takes settled agent loops the session has moved past and packs each into a short summary that holds intents, file paths, identifiers, errors, and code, landing at roughly 9 percent of the original tokens. The skeleton, meaning the system prompt, user messages, final answers, and paired tool calls with their results, is preserved byte for byte and never rewritten. Only the aged interior of past loops is stripped or packed, so the working set the model is reasoning over stays raw. ![Condense.chat](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/05-ceilings.jpg) On one real session replayed to 938 turns, Condense reports the bill falling 72.3 percent, taking a Sonnet run from 154 dollars to 43 and an Opus run from 771 to 214, with dollar-weighted savings across every session size near 66 percent. The company states the effect compounds with depth, since a larger context means a larger frozen prefix and a smaller slice of history re-read each turn, passing 53 percent on sessions with chains above 400,000 tokens. Answer faithfulness against uncompressed transcripts is reported at 94.2 percent. ![Condense.chat](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/07-zones.jpg) Setup is a single command that drops the harness in front of an existing agent with no key swap, and the drop-in API speaks both the OpenAI and Anthropic SDKs, so a team can point its base URL at a provider route and keep its own key. The proxy currently drives Claude Code, Codex, and OpenCode across macOS, Linux, and Windows. It is aimed at developers running agent workflows at scale, where re-read costs dominate, and the deepest sessions incur the largest bills. SPONSORED Condense is giving TestingCatalog readers 100M saved tokens. [Test it out! ](https://condense.chat/blog/anatomy-of-savings/?ref=testingcatalog.com) Condense is built by engineers with backgrounds shipping large-model infrastructure at Nord Security, Kilo Health, and Nexos AI. The full measurement pipeline is published as an open harness on GitHub, including the cost-split study and the replay tool, so the figures can be recomputed without a key. ### Vellum adds agent-to-agent AI collaboration for Slack URL: https://www.testingcatalog.com/vellum-adds-agent-to-agent-ai-collaboration-for-slack/ Last updated: 2026-07-02T23:21:18.000Z Vellum has launched agent-to-agent communication inside Slack, enabling individual AI assistants to coordinate work with one another and with the people they represent. This approach differs from a single shared assistant that an entire team interacts with. On Vellum, each person operates their own assistant, which holds its own memory, context, and permissions. These separate assistants can now converse in a common Slack channel and move tasks forward between them. > Today, we launched agent-to-agent conversations in Slack to give you real AI coworkers. > > Vellum assistants now talk to each other and coordinate work with your team all inside your workspace. > > We tested it with two agents in our own Slack. > > They planned our offsite for 19 people… [pic.twitter.com/HEZE5fAHKe](https://t.co/HEZE5fAHKe?ref=testingcatalog.com) > > — Marina · vellum.ai 👾 (@marinatrajk) [July 2, 2026](https://x.com/marinatrajk/status/2072748613051068504?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The launch demonstration focused on planning a company offsite. Two assistants, one belonging to a team member named Marina and another to a colleague named Akash, took the task into a shared channel and collaborated on it. They divided responsibilities, negotiated candidate dates, and then one assistant reached out to the wider team for venue and dietary preferences before finalizing the arrangements. The sequence demonstrated assistants communicating with each other and with humans within the same thread, rather than a single system handling every request. This capability is aimed at teams that already coordinate through Slack and want assistants that carry personal context into shared work. Since each assistant holds one person's history, preferences, and relationships, it can act as that person's representative when a plan requires input from several people. Permissions remain isolated by default, so one assistant accesses another's information only when necessary, and routine planning does not require anyone to restate context that an assistant already holds. Multi-step coordination, such as scheduling, collecting input, and confirming details, shifts from manual back-and-forth to assistants that manage the mechanics and present decisions to the people involved. SPONSORED Check agent-to-agent AI on Vellum! [Test it! ](https://vellum.ai/?utm%5Fsource=twitter&utm%5Fmedium=social&utm%5Fcampaign=feature-launch-07-02) Vellum Labs, the company behind the assistant, has built its product around personal, on-device intelligence and aims to give people back the hours spent on daily logistics. Its assistant is positioned as software that runs on a user's own machine and acts on their behalf, drafting messages, managing calendars, and handling tasks while keeping data under the user's control. The Slack feature extends that personal-assistant model outward, connecting one user's assistant to others so a team of individually informed assistants can operate together. This marks a transition from a single-user helper to shared workflows, in which each participant maintains a distinct assistant rather than pooling everything into a single account. ### Early look at Anthropic's Claude Science app for researchers URL: https://www.testingcatalog.com/early-look-at-anthropics-claude-science-app-for-researchers/ Last updated: 2026-07-01T22:28:54.000Z Anthropic has unveiled Claude Science ([Operon](https://www.testingcatalog.com/anthropic-tests-claude-operon-for-scientific-research-in-biology/)), an AI-powered research workbench designed specifically for scientists in the life sciences and related fields. This platform is currently available in beta for users on the Claude Pro, Max, Team, and Enterprise plans, supporting both macOS and Linux systems. > Introducing Claude Science, a new app designed with every stage of research in mind. > > Artifacts traced to their code, environments managed on demand, and 60+ optional scientific databases that you can connect. > > Available now in beta. [pic.twitter.com/HKhLknxLJO](https://t.co/HKhLknxLJO?ref=testingcatalog.com) > > — Claude (@claudeai) [June 30, 2026](https://x.com/claudeai/status/2072002740830842899?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Researchers can use Claude Science on their local machines, remote clusters, or via SSH, and it supports integration with existing high-performance computing infrastructure. The release targets scientific labs, academic institutions, and nonprofit organizations, with a special Team plan offering discounted rates for eligible research groups. Applications for dedicated AI for Science project credits are open until July 15, 2026, supporting up to 50 projects with substantial compute and platform credits. > ANTHROPIC 🔥: Early look at Claude Science, a new app designed for scientific research. > > \> "Artifacts traced to their code, environments managed on demand, and 60+ optional scientific databases that you can connect." > > Claude Science was named "Operon" during development and is… [pic.twitter.com/ExHnwRi5SW](https://t.co/ExHnwRi5SW?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 1, 2026](https://x.com/testingcatalog/status/2072338885724525007?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Claude Science brings together over 60 specialized tools and connectors for genomics, proteomics, cheminformatics, structural biology, and more. It manages complex analyses, handles computing job submissions, and ensures reproducibility by attaching code, environment details, and plain-language descriptions to every artifact. Built-in reviewer agents catch citation errors and calculation issues, while users can customize workflows or add proprietary lab tools as needed. Early users have leveraged the system for tasks spanning CRISPR screen design to multi-agent literature review, reporting dramatic reductions in analysis time and improvements in data validation. Anthropic, the developer, continues to gather feedback from research teams to refine platform capabilities and expand its reach in accelerating scientific discovery. [Source](https://www.anthropic.com/news/claude-science-ai-workbench?ref=testingcatalog.com) ### Google launches Nano Banana 2 Lite and Gemini Omni Flash URL: https://www.testingcatalog.com/google-launches-nano-banana-2-lite-and-gemini-omni-flash/ Last updated: 2026-07-01T22:28:42.000Z Google has announced the release of Nano Banana 2 Lite and Gemini Omni Flash, targeting developers and creators focused on multimedia generation and editing. Nano Banana 2 Lite is now available via the Gemini API and is recommended for users of the earlier Nano Banana model. It is built for environments where speed and budget are priorities, generating images from text prompts in just four seconds and costing $0.034 per 1,000 images. The model handles prompt adherence, character consistency, and text rendering with high reliability, and can be swapped in for immediate performance improvements over its predecessor. Nano Banana 2 Lite is also being integrated into Google consumer platforms, including Search, the Gemini app, and Google Photos. > Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio! > > Nano Banana 2 Lite is extremely fast (<4s image) & cheap ($0.034 / 1K image). > > Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast! [pic.twitter.com/qDxRpqpX5E](https://t.co/qDxRpqpX5E?ref=testingcatalog.com) > > — Logan Kilpatrick (@OfficialLoganK) [June 30, 2026](https://x.com/OfficialLoganK/status/2071988351083921690?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Gemini Omni Flash is now accessible in a public preview via Google AI Studio and the Gemini API. It allows developers to generate and edit up to ten seconds of video using multimodal inputs, including text, images, and short video clips. Video editing is conversational and supports multimodal referencing, enabling creators to maintain scene consistency and synchronize text or graphics with video actions. The model is priced at $0.10 per second of video. Early feedback from industry partners highlights the model’s ability to support rapid creative workflows and its promise for building advanced digital experiences. Limitations at launch include a 10-second cap on generation, a lack of audio input support, and some scene-consistency challenges. > We’re shipping 2 major releases:⁰ > 🔘 Nano Banana 2 Lite: our fastest and cheapest Gemini Image model > 🔘 Gemini Omni Flash: now available via the Gemini API and in [@GoogleAIStudio](https://x.com/GoogleAIStudio?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) to help developers generate and edit high-quality videos. [pic.twitter.com/fqB2sA5Xyl](https://t.co/fqB2sA5Xyl?ref=testingcatalog.com) > > — Google DeepMind (@GoogleDeepMind) [June 30, 2026](https://x.com/GoogleDeepMind/status/2071988044878516466?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Both models use SynthID watermarking for content verification. Google continues to build out its generative AI ecosystem, aiming to provide secure, scalable tools for developers and end users across multiple platforms. [Source](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/?ref=testingcatalog.com) ### Google might be testing Gemini Flash upgrade on LM Arena URL: https://www.testingcatalog.com/google-might-be-testing-gemini-flash-upgrade-on-lm-arena/ Last updated: 2026-07-01T14:08:41.000Z A Gemini Flash checkpoint appears to be circulating on [LM Arena](https://arena.ai/text?ref=testingcatalog.com), and early impressions place it a step above the Flash version currently running in the Gemini app. The gap looks incremental rather than generational, but testers comparing outputs are picking up a real difference in quality. Google hasn't commented on the listing, and it isn't clear whether this reflects a genuine release candidate or another internal build that quietly disappears. Google's Arena appearances have reliably preceded confirmed launches over the past year, which is part of why this one is drawing attention. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/Screenshot_20260701_074901_Chrome.jpg) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/svgviewer-png-output_47.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/svgviewer-png-output_49.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/07/svgviewer-png-output_53.webp) gemini-3-flash and gemini-3.5-flash responses on LM Arena The logical next step after Gemini 3.5 Flash would be a "**3.6**" label, but that's an extrapolation from Google's versioning habits rather than anything confirmed, and there's no indication of when or whether a wider rollout will follow. Another possibility could be a "**Gemini 4 Flash**" label, as its trace has already been spotted on GitHub. > GOOGLE 🔥: A new Gemini Flash checkpoint is being tested on LM Arena and may be released under a different version number. > > Gemini 3.6 Flash and even Gemini 4 Flash are among the possible options. > > Soon? 👀 [pic.twitter.com/ZrYsqnQMLe](https://t.co/ZrYsqnQMLe?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [July 1, 2026](https://x.com/testingcatalog/status/2072321107365957837?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) If a build like this surfaces officially, expect it first in the Gemini app's model picker, AI Studio, and the Gemini API, mirroring how the current Flash generation rolled out. Everyday Gemini users and cost-conscious developers stand to gain the most, since Flash carries the bulk of free and pay-as-you-go traffic that would otherwise need a pricier Pro-tier model. [Join Dev Mode ](https://discord.com/invite/devmode?ref=testingcatalog.com) That backdrop matters more than usual right now. Gemini 3.5 Flash launched at I/O in May as the default across the Gemini app and AI Mode in Search, beating the previous 3.1 Pro tier on several coding and agentic benchmarks while running several times faster. Gemini 3.5 Pro, pitched onstage for a June arrival, has since slipped into July, with reports pointing to additional tuning to coding, token efficiency, and long-task performance following early tester feedback. Whether that traces to Pro needing polish or to Google wanting more distance from OpenAI and Anthropic's coding benchmarks isn't confirmed, but rivals have kept pace on the agentic tasks Google has prioritized. Against that, a sharper Flash tier would yield a faster win, since Flash already carries most of the daily load for Google's fast-growing user base, while Pro remains unsettled. [Source](https://discord.gg/DevMode?ref=testingcatalog.com) ### Anthropic may impose KYC restrictions for Fable 5 access URL: https://www.testingcatalog.com/anthropic-may-impose-kyc-restrictions-for-fable-5-access/ Last updated: 2026-07-01T04:48:21.000Z UPDATE from 01.06.26: Anthropic will be restoring access to Claude Fable 5 globally for all paid users on Wednesday! > Claude Fable 5 will be available again globally tomorrow. > > After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding… > > — Anthropic (@AnthropicAI) [July 1, 2026](https://x.com/AnthropicAI/status/2072163884430229756?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Anthropic looks set to bring back Claude Fable 5, its most capable widely released model, and the latest build of the Claude app for iOS offers the clearest sign yet of how that return might be structured. Fresh strings referencing Fable 5 have appeared in the app, surfacing after the model was pulled from every surface earlier this month under a US government export control directive that disabled it globally. > ANTHROPIC 🔥: Claude Fable 5 is being prepared to run on usage credits that would also require identity verification. Sonnet 5 is being prepared for the release as well. > > \> Your credits will be added once your identity is verified. > > \> Fable 5 runs on usage credits, billed… [https://t.co/ibGRDsUGZQ](https://t.co/ibGRDsUGZQ?ref=testingcatalog.com) [pic.twitter.com/QuiAPgJ6Eg](https://t.co/QuiAPgJ6Eg?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 30, 2026](https://x.com/testingcatalog/status/2071923897520288164?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The wording points to two linked conditions: 1. Fable 5 may return only as a credit-gated option rather than as a bundled plan feature. This would align with the path Anthropic set at launch, when the model ran free on Pro, Max, Team, and Enterprise tiers from June 9 through June 22 before moving to paid usage credits. Because the suspension cut that free window short, subscribers who never got the full run of bundled access may find no fresh grace period waiting once it returns. 2. Identity verification — a possibility already raised within the developer community, and one the iOS strings appear to reinforce, with a know-your-customer step potentially required to receive Fable 5 credits. That verification layer is where the wider picture comes into focus. A know-your-customer gate would hand Anthropic a way to honor geographic limits tied to the export directive, most plausibly a US-only restoration. Read alongside the model's existing US-only inference option and its built-in block on distillation queries, the approach looks aimed at keeping frontier capability from being replicated abroad while restoring access for vetted domestic users. Timing stays unsettled. Anthropic executives have signaled a return within days, but no firm date or mechanics have been confirmed — though shipping-ready strings suggest an announcement is not far off. > 🚨 NEWS: Commerce is expected to lift export controls on Fable tonight, a senior White House official tells me. > > — Sophia Cai (@SophiaCai99) [June 30, 2026](https://x.com/SophiaCai99/status/2072082423949525348?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) UPDATE: Export Controls on Claude Fable 5 may be lifted as early as today, according to Politico. Join DevMode Discord for more! [Join it ](https://discord.com/invite/devmode?ref=testingcatalog.com) ### Anthropic launches Claude Sonnet 5 model on Claude and APIs URL: https://www.testingcatalog.com/anthropic-launches-claude-sonnet-5-model-on-claude-and-apis/ Last updated: 2026-06-30T18:33:59.000Z Anthropic has released Claude Sonnet 5, a new AI model designed to deliver advanced autonomous capabilities for developers and businesses. The model is available immediately for all users across Free, Pro, Max, Team, and Enterprise plans, as well as in Claude Code and on the Claude Platform. Pricing starts at $2 per million input tokens and $10 per million output tokens, with an increase scheduled after August 31, 2026\. Sonnet 5 can be accessed via the Claude API, enabling developers to integrate the model into their workflows and applications worldwide. > Introducing Claude Sonnet 5, our most agentic Sonnet yet. > > It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models. [pic.twitter.com/UKK8G7ww5h](https://t.co/UKK8G7ww5h?ref=testingcatalog.com) > > — Claude (@claudeai) [June 30, 2026](https://x.com/claudeai/status/2072017450611142835?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Claude Sonnet 5 boasts substantial improvements in reasoning, tool use, coding, and knowledge work compared to its predecessor, Sonnet 4.6\. Its agentic features enable it to plan and execute multi-step tasks, utilize web browsers and terminals autonomously, and complete projects that previously required more expensive models. Early access partners report that it performs reliably in complex technical tasks, follows through on multi-step assignments, and consistently refuses unsafe requests. > Claude Sonnet 5 is now available in Cursor. > > On CursorBench, it's a meaningful step up from Sonnet 4.6: 57% vs. 49%. [pic.twitter.com/AQVHzrvqcR](https://t.co/AQVHzrvqcR?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [June 30, 2026](https://x.com/cursor%5Fai/status/2072020786181988418?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Benchmarking data show that Sonnet 5 is closing the performance gap with the higher-end Opus 4.8 model at a significantly lower cost, making it particularly attractive for organizations seeking a balance between price and advanced capabilities. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/9941d610909f28a504e16dd5af823df172ec6035-2600x1234.webp) [Anthropic](https://www.testingcatalog.com/tag/claude/), the company behind Claude Sonnet 5, has focused this release on improving safety and reliability, especially in agentic contexts. The model scored lower on undesirable behaviors compared to earlier Sonnet versions and is deployed with real-time cyber safeguards similar to those used in Opus models. The release reflects Anthropic's ongoing commitment to safer, more capable generative AI tools for a broad developer and enterprise audience. [Source](https://www.anthropic.com/news/claude-sonnet-5?ref=testingcatalog.com) ### NoimosAI launches Creative Agent for brand assets URL: https://www.testingcatalog.com/noimosai-launches-creative-agent-for-brand-assets/ Last updated: 2026-06-30T16:05:02.000Z NoimosAI has shipped a Creative Agent that produces brand-ready assets from market evidence rather than blank-page prompts. The tool studies competitor output, top-performing creatives across the wider market, and a brand's own past results, then assembles assets built around patterns that are already converting. The pitch is direct: instead of guessing what a campaign should look like, teams start from what is measurably working and adapt it to their own brand. > NoimosAI can now turn market insights into high-performing content for your brand. > > It analyzes competitors, top creatives, and your past results to create assets grounded in what is already working. [pic.twitter.com/bZ9wAAIVXT](https://t.co/bZ9wAAIVXT?ref=testingcatalog.com) > > — NoimosAI (@noimos\_ai) [June 30, 2026](https://x.com/noimos%5Fai/status/2071987132986699817?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The agent scans high-performing content across platforms, including Meta, TikTok, and LinkedIn, isolates the structural patterns behind high-performing creatives, and maps those patterns onto a brand's voice, product, and prior campaigns. From there, output gets refined through conversation, inline comments, and direct edits, so a draft can be shaped before anything goes live. Finished assets download or publish straight to connected apps, keeping the path from idea to posted creative inside one workflow. ![NoimosAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-28-141100.png) That framing lands at a moment when producing creatives has stopped being the bottleneck for most marketing teams and distribution has taken its place. Generating assets is fast; generating assets that match proven winners and stay on-brand is harder, and that is the gap the Creative Agent targets. Availability is now open via a free trial on the NoimosAI site, with output routed to whichever connected channels a brand already uses. SPONSORED Start testing NoimosAI! [Learn more ](https://noimosai.com/en?ref=testingcatalog.com) NoimosAI positions the Creative Agent within a broader platform it describes as an all-in-one autonomous AI marketing team, one that connects to a brand's apps or website and runs work across social, SEO, outreach, and growth strategy on an ongoing basis. Built by AGOS LABS, the system is framed around autonomous execution grounded in a brand's own data and live market signals rather than generic templates or preset automation. The Creative Agent extends that approach into the creative layer, treating asset production as another function the platform runs and refines rather than a separate tool a team operates by hand. ### Apify lets AI Agents pay via Coinbase x402 for web tools URL: https://www.testingcatalog.com/apify-lets-ai-agents-pay-via-coinbase-x402-for-web-tools/ Last updated: 2026-06-30T15:14:07.000Z Apify has partnered with Coinbase to enable autonomous AI agents to discover, pay for, and run its web automation Actors without needing an Apify account, subscription, or API key. This integration utilizes x402, the payment protocol Coinbase formalized from the long-reserved HTTP 402 Payment Required status code. It addresses a significant gap in agent autonomy: while an agent could plan a task and call an API, it previously couldn't sign up for or pay for a service, necessitating human intervention. With x402, agents can independently settle payments, allowing work to proceed seamlessly. > Until today, agents could buy about 2,000 tools through x402. > > We just 10x'd that to 20,000+ 🚀 > > In partnership with [@coinbase](https://x.com/coinbase?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), we’re launching x402 support to give autonomous agents access to the largest marketplace of web automation tools. > > No account, API keys, or human in the… [pic.twitter.com/5OHLn6fvcI](https://t.co/5OHLn6fvcI?ref=testingcatalog.com) > > — Apify (@apify) [June 30, 2026](https://x.com/apify/status/2071963996576743795?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The process is straightforward. An agent sends a request to an Actor, receives an HTTP 402 response, authorizes a USDC micropayment from a wallet on the Base network, and the Actor executes the task. Settlement uses a stablecoin, ensuring one dollar remains equivalent to one dollar without volatility. Billing is pay-per-result, meaning the agent purchases a prepaid token and pays only for the usage they consume, avoiding the commitment to a plan. The most efficient method pairs the Coinbase Agentic Wallet CLI with a direct HTTP call to the Actor endpoint, in which the agent handles the 402 response. The same Actors are also accessible through the Apify MCP server and a client that automatically provisions a wallet. 0:00 /0:41 1× At launch, around 20,000 of Apify's Pay Per Event Actors are callable through x402, increasing the number of tools on the protocol from approximately 2,000 to over 20,000\. For agents, the cost of real work remains low. About one dollar can cover nearly 350 Google Maps listings, 430 Instagram profiles, or 110 trending TikTok videos in a single run. Since x402 is an open protocol governed by the Linux Foundation, the same method can be applied across the broader agentic economy rather than being limited to a single vendor endpoint. Apify Actors are cloud programs designed for web scraping and browser automation, covering search, profile collection, and structured data extraction across major platforms. Most x402 services to date have been single, purpose-built endpoints, so opening a community-driven marketplace of this scale to agents in one step provides the ecosystem with a long tail of ready tools it previously lacked. The target audience includes teams building autonomous agents, developer tools, and internal copilots that need to access live web data and act on it without requiring human approval at each step. SPONSORED Check out their blogpost!! [Learn more ](http://blog.apify.com/introducing-x402-agentic-payments/?utm%5Fsource=testingcatalog&utm%5Fmedium=social&utm%5Fcampaign=gtm-cam-068) Apify operates a platform and marketplace for web automation, where developers can publish and monetize Actors that other builders, and now agents, can run on demand. The company positions itself as the first and largest marketplace of web automation tools for agents on x402\. Coinbase provides the protocol layer, transforming a status code that had been unused since the early days of the web into a functional standard for machine-to-machine payments. Together, this initiative places a vast catalog of web tools within the pay-per-call model that autonomous agents need to operate end-to-end. ### Bloome launches chat platform for AI agent teams URL: https://www.testingcatalog.com/bloome-launches-chat-platform-for-ai-agent-teams/ Last updated: 2026-06-30T13:41:29.000Z Bloome has launched an instant messaging platform built around human-agent teams, where people and AI agents work within a single shared chat. Instead of a single model answering alone, several agents take part in the same conversation: one drafts a response, another pushes back on it, and another catches what is missing, so the version that survives the exchange is the strongest. Bloome brings leading models together in a single place, including Claude, ChatGPT, Gemini, Grok, and DeepSeek, and it connects coding agents such as Claude Code, Codex, Gemini CLI, and OpenCode, with the option to add custom agents through a one-click connection. Every contribution from teammates and agents is consolidated into a single document that provides the full context, and anyone with access can revisit it later. 0:00 /0:16 1× The platform aims at non-technical knowledge workers, the people who move work forward through communication, coordination, and decision-making, including product managers, marketers, operators, analysts, consultants, and founders. Bloome frames a set of practical workflows around this group, from market research and code review to data analysis, contract review, presentation building, and creative campaign development. In a research workflow, for example, one agent drafts findings, a reviewer agent challenges the assumptions and sources, the agents refine the report together, and the team receives a final result that it can rely on. The work lands in shared outputs such as project dashboards and campaign assets, and Bloome is available for download on macOS, Windows, iOS, and Android, where new members joining a conversation can be brought up to speed on prior context. SPONSORED Test it out with Bloome! [Learn more ](https://bloome.im/?ref=testingcatalog.com) The release arrives as multi-agent collaboration becomes a focus across the AI tooling space, with builders moving past single-model chat toward setups where several agents and several people reason over the same material at once. Bloome positions its answer around a network for human-agent teams rather than a solo assistant, putting cross-checking among agents, shared memory, and team-ready outputs in a single workspace. The product carries the line that workers are amplified rather than replaced, and the team behind it is rolling out the launch to a broad audience of knowledge workers who want more than a single model's first answer for the work that matters to them. Bloome is now reachable through its website and across desktop and mobile apps, with an agent marketplace and a skill marketplace that extend what teams can plug in. ### Meituan launches LongCat-2.0 1.6T parameter model on APIs URL: https://www.testingcatalog.com/meituan-launches-longcat-2-0-1-6t-parameter-model-on-apis/ Last updated: 2026-06-30T11:00:35.000Z Meituan has unveiled LongCat-2.0, marking a significant advancement in its LongCat model family following the earlier LongCat-2.0-Preview. This new model is designed as a 1.6 trillion-parameter Mixture-of-Experts system, with approximately 48 billion parameters active per token. It is aimed at agentic coding, tool use, long-context work, automated workflows, and the execution of complex instructions. LongCat-2.0 features a 1 million-token context window and a maximum output length of 128K tokens via the LongCat API Platform. Developers can access it through OpenAI-compatible and Anthropic-compatible API formats, with support for Claude Code, OpenClaw, OpenCode, Kilo Code, and Codex-style workflows. > Introducing LongCat-2.0 🐱 > 1.6T parameters · MoE with \~48B active · 1M context > The full model behind Owl Alpha on [@OpenRouter](https://x.com/OpenRouter?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) — now available. > > Built for agentic coding from the ground up: > ◆ LongCat Sparse Attention (LSA) — scales efficiently for 1M-context tokens > ◆… [pic.twitter.com/zum2SdZ0Z2](https://t.co/zum2SdZ0Z2?ref=testingcatalog.com) > > — Meituan LongCat (@Meituan\_LongCat) [June 30, 2026](https://x.com/Meituan%5FLongCat/status/2071783587205308721?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The company reports that the full training run and deployment were conducted on AI ASIC superpods, with pretraining across more than 35 trillion tokens. LongCat also introduced LongCat Sparse Attention for long-horizon tasks and trained the model on hundreds of billions of tokens of 1M-context data, positioning the system for large repositories, long documents, and multi-step agent tasks. The release is publicly available via the API, and billing is now active. The pay-as-you-go pricing structure currently supports LongCat-2.0 at: 1. $0.75 per 1M uncached input tokens 2. $0.015 per 1M cached input tokens 3. $2.95 per 1M output tokens Lower limited-time prices are also listed by LongCat. Token packs are valid for 30 days, and cache hits do not count against token-pack usage. This release is not yet a full weights drop. The GitHub repository is public under an MIT license, but both the repository and Hugging Face model card indicate that model weights are forthcoming. This makes the launch a hybrid release for now: usable through the API and documented in public repositories, while the downloadable model weights remain pending. > Some of you guessed right. 👀 > Owl Alpha on [@OpenRouter](https://x.com/OpenRouter?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) — that's us. > > Since going live, it has reached Top 3 globally by daily volume — and #1 on Hermes Agent, #2 on Claude Code, #3 on OpenClaw by monthly volume. > > Thank you to everyone who tested and used Owl Alpha during stealth… [pic.twitter.com/e86L9x3hFI](https://t.co/e86L9x3hFI?ref=testingcatalog.com) > > — Meituan LongCat (@Meituan\_LongCat) [June 29, 2026](https://x.com/Meituan%5FLongCat/status/2071624742701080606?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) LongCat-2.0 is also linked to Owl Alpha, the previously undisclosed model running on OpenRouter. LongCat’s official account describes LongCat-2.0 as the full model behind Owl Alpha, while OpenRouter lists Owl Alpha as a 1.05M-context agentic model with tool-use, code-generation, automated workflow, and complex instruction-following capabilities. OpenRouter’s free-models page lists Owl Alpha at 3.74T tokens, indicating the model had already seen significant developer usage before the reveal. Meituan, the company behind LongCat, describes the project as a family of large language models designed to make AI useful in physical-world scenarios. The team has already released LongCat-Flash-Chat, LongCat-Video, LongCat-Image, LongCat-Next, and other AI projects, positioning LongCat-2.0 as the new flagship language model in a broader multimodal and agent-focused portfolio. [Source](https://longcat.chat/blog/longcat-2.0/?ref=testingcatalog.com) ### Cursor releases its iOS app for vibe coding on the go URL: https://www.testingcatalog.com/cursor-releases-its-ios-app-for-vibe-coding-on-the-go/ Last updated: 2026-06-29T22:34:16.000Z Cursor has released a [native iOS app](https://apps.apple.com/us/app/cursor/id6767085653?ref=testingcatalog.com) that transforms the phone into a control surface for its coding agents. The app is available in public beta for users on paid Cursor plans, with App Store distribution for iPhone and a required iOS version of 26.0 or later. The download itself is free with in-app purchases, while Cursor states that beta access is tied to paid plans. The app allows developers to: 1. Choose a repository 2. Launch agents 3. Select frontier models 4. Use voice input 5. Issue slash commands 6. Review work 7. Inspect diffs 8. Leave follow-up instructions 9. Merge pull requests from a phone It also supports Remote Control for agents running on a local computer, including a setting that keeps the machine awake so the session remains reachable away from the desk. > Introducing Cursor for iOS. > > Build from anywhere by launching always-on cloud agents. Or remotely control agents running on your computer from the app. > > Composer 2.5 is 75% off in the app now through July 5\. [pic.twitter.com/dFxQyrgmBb](https://t.co/dFxQyrgmBb?ref=testingcatalog.com) > > — Cursor (@cursor\_ai) [June 29, 2026](https://x.com/cursor%5Fai/status/2071641103191998810?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Cloud agents run in isolated virtual machines with full development environments, where they can test, verify, generate demos, screenshots, and logs, and then hand off work to the developer for review. Local sessions can also be moved to the cloud, so work can continue without keeping a laptop open. The App Store listing provides more product details: users can annotate screenshots, review videos and logs, talk to agents via voice, and use models from OpenAI, Anthropic, Google, and others, as well as Cursor’s own Composer model. Cursor is also offering 75% off Composer 2.5 runs in the mobile app through July 5, 2026. Anysphere, Inc., the company behind Cursor, positions itself as an applied research lab focused on the future of programming. Cursor’s broader product line already spans agents, cloud agents, CLI, code review, Tab, enterprise tooling, and marketplace integrations, making the iOS app part of a wider push to make agentic development less tied to a single desktop session. [Source](https://cursor.com/blog/ios-mobile-app?ref=testingcatalog.com) ### OpenAI prepares upgraded Excel and PowerPoint controls for Codex URL: https://www.testingcatalog.com/openai-prepares-upgraded-excel-and-powerpoint-controls-for-codex/ Last updated: 2026-06-28T08:27:54.000Z OpenAI appears to be preparing a more targeted form of computer use for Codex, with two options surfacing in the [Computer Use](https://www.testingcatalog.com/openai-codex-transformed-into-superapp-with-computer-use/) section of the settings, both tied to using Microsoft Office. Rather than treating PowerPoint and Excel as generic windows to click through, the toggles suggest Codex working through each app's add-in layer, the task-pane and content add-in framework Microsoft exposes to developers, to handle slides and spreadsheets with more precision than screenshot-and-cursor control allows. Codex's computer use, shipped earlier this year, already reads native interfaces and drives them through clicks and keystrokes, but that general approach struggles with the depth of a pivot table or the master-slide structure of a deck. Routing through add-ins would give Codex a structured handle on document internals, the kind of access that separates a reliable edit from a brittle one. For people who live in spreadsheets and decks, finance teams, analysts, and consultants, it would cut the shuffle between Codex and Office and let more work be handed off in place. The features remain unreleased and carry no public timeline; they sit in settings without being switched on. Where they would appear is clear enough: inside the desktop Codex app on macOS and Windows, as opt-in controls beside the existing Computer Use plugin. The strategic read is the more telling part. [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) has spent the year turning Codex from a coding assistant into a general workbench that operates the software around it, and a dedicated Office path fits that arc. It also sets Codex against Microsoft's own Copilot, whose agent mode now supports multi-step actions natively in Word, Excel, and PowerPoint, and which relies on Anthropic's models for some of that work. Anthropic, in turn, drives Office through its own add-ins. Operating the productivity layer through structured hooks rather than raw pixels is becoming the shared battleground, and Codex is moving to claim its share. ### OpenAI tests gifting Codex credits as new growth strategy URL: https://www.testingcatalog.com/openai-tests-gifting-codex-credits-as-new-growth-strategy/ Last updated: 2026-06-28T08:23:04.000Z OpenAI appears to be preparing a gifting layer for Codex that would let people pass credits to others rather than keeping usage tied to a single account. A recent build surfaces a hidden widget designed to help friends bring their ideas to life by sharing credits, plus a Gifts entry on the user profile page. The link points to a page that is not yet live, indicating groundwork ahead of a rollout rather than anything usable today. No timeline is attached, and the mechanics, such as how many credits can move, whether recipients must be new to Codex, and what caps or expirations apply, remain unsettled. ![Codex](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-25-at-22.33.58.png) The feature sits awkwardly with OpenAI's published terms, which currently state that Codex credits are non-transferable and not giftable. That implies any gifting flow would revise those terms or route through a separate promotional structure, closer to how referral rewards work than to a transfer of purchased balance. It would most plausibly appear in the same lower-left profile menu where invite-a-friend flows already live. > We heard you wanted to use Codex rate limit resets on your own time. > > Starting today, we’re rolling out the ability to save rate limit resets to use later. > > We’re starting Go, Plus, Pro, and Business users with one free reset: [pic.twitter.com/gucyTi04wc](https://t.co/gucyTi04wc?ref=testingcatalog.com) > > — OpenAI (@OpenAI) [June 12, 2026](https://x.com/OpenAI/status/2065225362544726371?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The timing is the tell. [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) ran a referral pilot from June 11 through June 24 in which eligible Plus and Pro users could invite up to three friends. Once a recipient sent their first Codex message, both people received a banked rate-limit reset. That window has now closed. Gifting credits reads as the next probe along the same line: a loop that pulls new and lapsed users into Codex through existing users, with a clear path toward selling bundles later. For a product competing with Anthropic's Claude Code and Cursor for daily developer attention, turning the user base into a distribution channel is cheaper than discounts, and a credit economy that can be gifted, earned, and eventually bought tends to grow once the plumbing is in place. Whether it graduates from a dead link to a shipping feature is the question worth watching. ### Microsoft launches MAI-Code-1-Flash on GitHub Copilot URL: https://www.testingcatalog.com/microsoft-launches-mai-code-1-flash-on-github-copilot/ Last updated: 2026-06-27T17:57:24.000Z Microsoft has launched MAI-Code-1-Flash, its proprietary AI coding model, to a broad user base through general availability for GitHub Copilot Business and Enterprise subscribers. The model is now accessible to organizations on these plans, provided administrators activate the relevant policy in Copilot settings. MAI-Code-1-Flash is engineered for rapid, low-latency code generation, catering to professional developers and large teams who rely on fast, iterative coding cycles in complex software projects. The model is priced according to provider list rates within usage-based billing, aligning with Copilot’s overall pricing structure. > .[@MicrosoftAI](https://x.com/MicrosoftAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com)'s MAI-Code-1-Flash is now generally available for GitHub Copilot Business and Copilot Enterprise. > > Built for coding and optimized for GitHub Copilot, MAI-Code-1-Flash delivers fast, low-latency responses that make it well-suited for high-volume, iterative agentic… > > — GitHub Enterprise (@GitHubEnt) [June 26, 2026](https://x.com/GitHubEnt/status/2070553293244321969?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) MAI-Code-1-Flash stands out due to its purpose-built architecture, focusing on high-speed code completion and agentic workflows that prioritize efficiency in large-scale development environments. Unlike earlier iterations or more general-purpose coding models, this release delivers notable performance improvements in both speed and response times, addressing a key demand among enterprise customers for scalable, responsive AI coding solutions. Early feedback from developer communities highlights its responsiveness and reliability in managing intensive coding workloads. Microsoft, as the parent company of GitHub and a major AI research entity, continues to invest in advanced AI-powered developer tools. This release underscores its strategy to provide enterprise-grade AI coding support, further integrating its in-house AI capabilities into widely used platforms like GitHub Copilot. The move strengthens Microsoft’s position in the growing market for AI-assisted software development tools. [Source](https://github.blog/changelog/2026-06-26-mai-code-1-flash-for-copilot-business-and-copilot-enterprise/?ref=testingcatalog.com) ### ICYMI: Google adds Computer Use to Gemini 3.5 Flash URL: https://www.testingcatalog.com/icymi-google-adds-computer-use-to-gemini-3-5-flash/ Last updated: 2026-06-27T17:48:46.000Z Google has announced that computer use is now integrated directly into Gemini 3.5 Flash, marking a major step forward in its Gemini AI platform. This development opens the door for developers and enterprises to leverage agentic computer use tasks with the model's highest performance to date. Previously, computer use capabilities were limited to a standalone Gemini 2.5 computer use model, but now they are available natively in the latest Gemini Flash version. The rollout is available to developers and enterprise users via the Gemini API and Gemini Enterprise Agent Platform, targeting organizations seeking advanced automation solutions. > Excited to introduce Computer Use support for Gemini 3.5 Flash!🔥 > > This enables Gemini to reason and act across platforms (browser, mobile, and desktop environments) > > We see significant improvements across many work-related automation tasks, from filing tickets and more. Enjoy! [pic.twitter.com/Yy3tGvHx0D](https://t.co/Yy3tGvHx0D?ref=testingcatalog.com) > > — Omar Sanseviero (@osanseviero) [June 24, 2026](https://x.com/osanseviero/status/2069821925128647148?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Gemini 3.5 Flash's built-in computer use enables custom agents to observe, reason, and act across browser, mobile, and desktop environments. This supports complex use cases such as continuous software testing and enterprise knowledge work. Technical upgrades include targeted adversarial training to address prompt injection risks and two optional enterprise safeguard systems: 1. Requiring explicit user confirmation for sensitive actions. 2. Halting tasks if indirect prompt injections are detected. Gemini’s approach encourages combining these measures with sandboxing, human-in-the-loop verification, and strict access controls. Early enterprise adopters have already reported value from deploying these features in live environments. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/gemini-3-5__benchmark-OSWorld-Ve.width-1000.format-webp--1-.webp) Google continues to expand Gemini’s capabilities, positioning the platform to meet evolving demands for AI-powered automation in business settings. This latest update demonstrates Google’s focus on security, flexibility, and multi-environment support for enterprise customers. [Source](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-computer-use-gemini-3-5-flash/?ref=testingcatalog.com) ### Google tests notebook collections for NotebookLM URL: https://www.testingcatalog.com/google-tests-notebook-collections-for-notebooklm/ Last updated: 2026-06-27T12:42:31.000Z Google appears to be developing a collections system for NotebookLM that would let people group several notebooks under a single heading, surfaced via a dedicated tab in the main navigation. The capability has not been officially confirmed and looks to be early in development, with no timeline attached, but it points to a gap the tool has carried for some time. NotebookLM already reads and clusters the sources inside an individual notebook, a source-level organization layer that reached full rollout in early May 2026\. What it has lacked is any native way to group whole notebooks. Power users have leaned on browser extensions to approximate folders, and the team itself has acknowledged notebook-level grouping as the main piece still missing. Collections would fill that role, providing a top-level structure rather than a flat, scrolling list. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/NotebookLM-06-27-2026_02_41_PM.jpg) The people most likely to benefit are heavy users juggling many notebooks, particularly those who treat notebooks as project workspaces within Gemini. The two products now sync, so a source added in one place appears in the other. Notebook projects went free for all Gemini web users earlier this year, so the population sitting on large libraries has grown quickly. With the free tier allowing up to 100 notebooks, navigation friction was always going to surface. There are signs that Google first weighed a label-based approach to notebook organization before leaning toward collections, though the reasoning behind any shift remains unclear. The move fits a year spent reshaping [NotebookLM](https://www.testingcatalog.com/tag/notebooklm/) from a question-and-answer layer over documents into a research-to-output hub welded to Gemini. Rivals face the same organizational weak spot: ChatGPT Projects, Claude Projects, and Perplexity Spaces all offer single containers, but none gracefully support grouping across dozens of them. A clean top-level layer would give Google an early answer to a problem the whole category is starting to feel. ### OpenAI launches GPT-5.6 Sol preview for select partners URL: https://www.testingcatalog.com/openai-launches-gpt-5-6-sol-preview-for-select-partners/ Last updated: 2026-06-26T21:16:40.000Z OpenAI is opening a limited preview of GPT-5.6, led by Sol, its new flagship model, alongside Terra for lower-cost everyday work and Luna for faster, cheaper workloads. The preview starts with a small group of trusted partners, with access initially through the API and Codex, while broader access for ChatGPT, Codex, and API users is planned in the coming weeks. > Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5\. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price. > > Bad news: at the request of the US government, it is launching today in… > > — Sam Altman (@sama) [June 26, 2026](https://x.com/sama/status/2070607488274358364?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) GPT-5.6 Sol arrives as OpenAI’s strongest model in the new family, with gains positioned around agentic coding, biology workflows, and cybersecurity tasks. The model adds a new max reasoning effort for deeper problem solving and an ultra mode that uses subagents to work on complex tasks beyond a single-agent setup. OpenAI says Sol sets a new state of the art on Terminal-Bench 2.1 and shows stronger GeneBench v1 results than GPT-5.5 while using fewer tokens. ![OpenAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/HLwWT8TagAALRWm.jpeg) The release is being handled as a controlled rollout because of the model’s cyber and biological capabilities. OpenAI says Sol, Terra, and Luna are classified as High capability in both Cybersecurity and Biological and Chemical risk under its Preparedness Framework, while not reaching the High threshold for AI self-improvement or the Cyber Critical threshold. In browser exploit tests involving Chromium and Firefox, Sol identified bugs and exploitation primitives but did not autonomously produce a full-chain exploit under the tested conditions. OpenAI is pairing the model upgrade with a layered safeguard stack covering: 1. Model-level refusal behavior 2. Real-time cyber and biology misuse classifiers 3. Account-level review 4. Differentiated access 5. Monitoring, enforcement, and ongoing testing Some preview users may see blocked requests or slower responses when generation is paused for extra review, especially in dual-use security contexts where defensive and offensive work can initially look similar. The company says it used more than 700,000 A100-equivalent GPU hours for automated red teaming focused on universal jailbreaks, alongside human expert and third-party testing. OpenAI plans to keep testing during the preview period and publish an updated system card when the GPT-5.6 family moves toward general availability. Pricing starts at $5 per 1 million input tokens and $30 per 1 million output tokens for Sol, $2.50 input and $15 output for Terra, and $1 input and $6 output for Luna. GPT-5.6 also adds explicit cache breakpoints, a 30-minute minimum cache life, cache writes at 1.25x the uncached input rate, and cache reads with a 90% cached-input discount. OpenAI also plans to launch GPT-5.6 Sol on Cerebras in July at up to 750 tokens per second for select customers. [Source](https://openai.com/index/previewing-gpt-5-6-sol/?ref=testingcatalog.com) ### Microsoft adds Copilot finance tools to Excel for M365 users URL: https://www.testingcatalog.com/microsoft-adds-copilot-finance-tools-to-excel-for-m365-users/ Last updated: 2026-06-26T21:13:27.000Z Microsoft is rolling out a suite of new capabilities for Copilot in Excel, targeting finance professionals such as FP&A teams, accountants, tax specialists, and treasury managers who rely on Excel for modeling, analysis, and reporting. These updates include specialized “skills” for automating repeatable financial workflows, expanded financial data connectors, and improved traceability features. The features are now generally available to Microsoft 365 Copilot customers on Excel for Web, Windows, and Mac, with custom skills accessible through the Insiders channel and broader rollout planned for the following month. Partner-built skills are expected in Q3 2026. > Today we’re bringing skills to Copilot for Excel, giving teams a new way to scale their expertise across every workbook. [pic.twitter.com/DfD1nfPEO3](https://t.co/DfD1nfPEO3?ref=testingcatalog.com) > > — Satya Nadella (@satyanadella) [June 25, 2026](https://x.com/satyanadella/status/2070180313654063255?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The new skills allow finance teams to define structured processes for tasks like DCF modeling, variance analysis, and monthly reporting, reducing repetitive setup and ensuring consistency and accuracy. Users can create custom skills via open-standard markdown files, which Copilot reads to automate tailored financial processes. New connectors integrate data from CB Insights, Daloopa, FactSet, Morningstar, PitchBook, and S&P Global, bringing relevant financial, market, and company data directly into Excel. These additions offer broader coverage and deeper integration than previous versions and rival offerings, especially for firms requiring comprehensive, up-to-date data within their financial models. Microsoft’s collaboration with the Financial Modeling Institute ensures that Copilot in Excel meets rigorous industry standards, evaluating features against real-world finance cases. Early feedback from Microsoft's internal finance teams and partners has influenced product refinement, with experts noting the value of traceable, reviewable changes and access to trusted data sources. These improvements reflect Microsoft's strategy to address the unique demands of finance workflows and reinforce its position in the productivity and enterprise software space. [Source](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/25/copilot-in-excel-built-for-the-era-of-frontier-finance/?ref=testingcatalog.com) ### DeepReinforce releases Ornith-1.0 open-source coding models URL: https://www.testingcatalog.com/deepreinforce-releases-ornith-1-0-open-source-coding-models/ Last updated: 2026-06-25T14:27:22.000Z DeepReinforce has open-sourced Ornith-1.0, a self-improving family of models built for agentic coding. The release spans the full range, from a compact 9B Dense version meant for edge deployment up to a 397B MoE model aimed at frontier-scale work, with 31B Dense and 35B MoE options in between. Each variant is trained on top of pretrained Gemma 4 and Qwen 3.5 foundations. ![Ornith-1.0](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/HLqjo1CbMAAoc8u.jpeg) Ornith-1.0 What sets Ornith-1.0 apart from most reinforcement learning setups is how it handles the scaffold. Rather than depending on human-designed harnesses to steer solution generation, the model learns to produce both the solution rollouts and the task-specific scaffolds that guide them. Each RL step runs in two stages. Conditioned on a task and the scaffold last used for it, the model proposes a refined scaffold, then generates a solution conditioned on that scaffold. Reward from the rollout flows back to both stages, so the model is trained to author the orchestration as well as the answer. Repeated across training, scaffolds get mutated and selected toward those that produce higher-reward trajectories, and per-task strategies surface on their own without hand-engineered harness design. Letting a model write its own scaffold opens a path to reward hacking, where a scaffold satisfies the verifier without doing the task. DeepReinforce describes a three-layer defense: 1. A fixed outer trust boundary that keeps the environment and test isolation beyond the model's reach. 2. A deterministic monitor that flags attempts to read withheld paths or alter verification scripts. 3. A frozen LLM judge that vetoes the verifier when gaming happens inside the allowed tool surface. ![Ornith-1.0](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/image-8.png) Ornith-1.0 On performance, DeepReinforce positions Ornith-1.0 as state of the art among open-source models of comparable size. The company reports the 397B flagship reaching 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified, figures it says match Claude Opus 4.7 and top open peers such as MiniMax M3 and DeepSeek-V4-Pro. The 35B model is reported to clear similarly sized Qwen and Gemma builds, while the 9B version is said to hit 43.1 on Terminal-Bench 2.1 and 69.4 on SWE-Bench Verified and match far larger models like Gemma 4-31B, which puts capable coding within reach of resource-limited hardware. SPONSORED Check the models out on HuggingFace! [Learn more ](https://huggingface.co/deepreinforce-ai?ref=testingcatalog.com) DeepReinforce is the AI lab behind the release, a team that publishes reinforcement learning research in the open, including prior work such as CUDA-L1, and that shipped the IterX optimization loop for code agents. Ornith-1.0 carries that direction further by folding scaffold construction into the training process itself. The weights and a technical report are released on Hugging Face for teams that want to run or study the models directly. ### Google tests voice dictation and Magic Pointer on Gemini desktop URL: https://www.testingcatalog.com/google-tests-voice-dictation-and-magic-pointer-on-gemini-desktop/ Last updated: 2026-06-24T16:11:16.000Z Google's [Gemini desktop](https://www.testingcatalog.com/exclusive-early-look-at-the-next-gemini-desktop-upgrade/) app for macOS is poised for a significant voice upgrade, with early indications suggesting three new features beyond the basics already available on mobile. The main **Gemini Live interface has been redesigned** to resemble the phone layout: a full-screen canvas with a glowing center and control buttons anchored at the bottom. This change signals Google's intention to create a unified voice interface across devices rather than maintaining two separate ones. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-24-at-00.46.04.webp) The first addition is **system-wide voice dictation**. An in-app prompt describes the ability to summon a Gemini panel that types for you in any application. You can assign a hotkey, switch to your browser or editor, and speak. In practice, it functions as a voice keyboard overlaying the entire machine, aligning with the screen-aware drafting feature Google showcased at I/O, where spoken thoughts are transformed into clean text wherever the cursor is positioned. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-22-at-23.02.32.webp) The second feature, reminiscent of the **Magic Pointer** concept [shown earlier](https://deepmind.google/blog/ai-pointer/?ref=testingcatalog.com), allows Gemini to follow whatever the cursor hovers over. This ensures that both the user and the model remain focused on the same on-screen element during a spoken interaction. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-24-at-00.45.56.webp) The third addition is less clear: a menu entry, located beside video and image generation options, for **connecting to other macOS devices**. Its purpose is still uncertain, though it suggests a potential path toward one desktop instance controlling another. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-24-at-00.45.19.webp) This trajectory aligns with Google's stated plan to introduce its [Gemini Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/) agent and enhanced voice capabilities to the Mac client this summer, bridging the gap between the desktop app and the web. It also positions Google alongside OpenAI and Anthropic, which already offer similar features such as [Codex Remote Control](https://www.testingcatalog.com/openai-will-let-codex-control-other-desktop-devices-via-computer-use/) and Claude's Dispatch. Currently, these features are being tested by a small group, so the final version that becomes widely available may still undergo changes. Join [DevMode Discord](https://discord.com/invite/devmode?ref=testingcatalog.com) for more 👀 ### Meta launches AI glasses with three new styles from $299 URL: https://www.testingcatalog.com/meta-launches-ai-glasses-with-three-new-styles-from-299/ Last updated: 2026-06-23T22:37:20.000Z Meta has announced the launch of Meta Glasses, a new line of AI-powered eyewear developed in collaboration with EssilorLuxottica. The release includes three distinct frame styles: Meta Adventurer, Meta Fury, and Meta Glasses by Kylie, the latter created in partnership with Kylie Jenner. These glasses are designed for a wide audience, offering 26 style variations and compatibility with prescription lenses. Starting at $299, Meta Glasses are now available for purchase in multiple countries through Meta.com, Best Buy, Amazon, Lenscrafters, Sunglasses Hut, and additional select retailers. > Meta announced a new series of Meta Glasses in partnership with EssilorLuxottica. > > \> Compatible with prescription lenses. > \> 26 styles across a range of colors, lenses, and frames. > \> Launching with Meta AI powered by Muse Spark from day one. > > My Meta HSTN still didn't get Muse… [pic.twitter.com/B1FoWsG96J](https://t.co/B1FoWsG96J?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 23, 2026](https://x.com/testingcatalog/status/2069537816674271363?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The frames are crafted from premium materials, and each pair is equipped with advanced features such as: 1. Three-way adjustable nose pads 2. Open-ear speakers for audio 3. A multi-mic array with wind noise reduction 4. Hands-free photo and video capture The dedicated action button allows users to quickly access Meta AI or customize shortcuts. Battery life supports over 8 hours of use, with the included charging case extending usage by up to 40 hours. Meta Glasses are powered by Meta AI running on the new [Muse Spark model](https://www.testingcatalog.com/meta-to-release-muse-spark-in-voice-mode-and-meta-glasses/), offering improved multimodal understanding and support for 14 additional live translation languages. Industry analysts note the expanded functionality and focus on privacy controls, as well as the integration of new features like dynamic photo and upcoming pedestrian navigation. [Meta](https://www.testingcatalog.com/tag/meta/), formerly Facebook, has been investing in wearable AI technology to create seamless daily digital experiences. This partnership with EssilorLuxottica leverages both companies’ experience in technology and eyewear design, aiming to position Meta Glasses at the forefront of AI-powered personal devices. [Source](https://about.fb.com/news/2026/06/meta-essilorluxottica-partner-launch-meta-glasses/?ref=testingcatalog.com) ### Anthropic launches Claude Tag on Team and Enterprise plans URL: https://www.testingcatalog.com/anthropic-launches-claude-tag-on-team-and-enterprise-plans/ Last updated: 2026-06-23T22:23:10.000Z Anthropic is moving Claude deeper into workplace chat with Claude Tag, a new Slack-based agent that teams can summon by typing @claude in a channel or thread. The beta is available now for Claude Enterprise and Team customers, and Anthropic says the rollout starts in Slack before expanding to other places where teams work. Claude Tag works with Opus 4.8 and replaces the older Claude in Slack app, with admins given 30 days to opt in before the prior Slack experience switches over on August 3, 2026. > Introducing Claude Tag, a new way for teams to work with Claude. > > In Slack, Claude joins as a team member with access to the channels and tools you choose. Tag Claude in and delegate tasks to it while you focus on other work. [pic.twitter.com/R2C6A5Kcye](https://t.co/R2C6A5Kcye?ref=testingcatalog.com) > > — Claude (@claudeai) [June 23, 2026](https://x.com/claudeai/status/2069468693017268244?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The feature turns Claude from a single-user assistant into a shared team agent. Once added to selected Slack channels, Claude can read thread context, break requests into steps, use connected tools and data, post results back into the same thread, schedule follow-ups, and continue work over hours or days. Anthropic says Claude can also remember relevant channel context, flag unresolved threads, monitor channels, pull metrics, prepare call briefings, triage backlogs, and open draft PRs from bug reports when connected to the right systems. The release is aimed at organizations rather than individual users. Admins decide which channels Claude can join, which tools and repositories it can access, and how much the organization or each channel can spend. Claude Tag runs under an “agent identity” model, meaning Claude acts through its own service accounts instead of borrowing one user’s login. Anthropic says this allows channel-scoped permissions, separate memory between private workspaces, audit logs for network calls and actions, and admin controls over who can use the feature. This also changes the billing model. Channel work is billed to the organization, while direct messages with Claude in Slack are billed to the individual user’s Claude account. Anthropic is offering a launch credit for eligible customers: 1. $25,000 for Claude Enterprise organizations 2. $2,500 for Claude Team organizations with at least 10 paid seats These credits expire on September 1, 2026. For [Anthropic](https://www.testingcatalog.com/tag/claude/), Claude Tag is positioned as an evolution of Claude Code and Claude Cowork into multiplayer AI work. The company says tagging @claude has become one of its own internal workflows, including code creation, product metrics, support tickets, and debugging. Anthropic claims that 65% of its product team’s code is created by its internal version of Claude Tag, making this launch both a customer release and a public test of its own AI-native operating model. [Source](https://www.anthropic.com/news/introducing-claude-tag?ref=testingcatalog.com) ### Mistral launches OCR 4 for multilingual document extraction URL: https://www.testingcatalog.com/mistral-launches-ocr-4-for-multilingual-document-extraction/ Last updated: 2026-06-23T22:03:51.000Z Mistral has announced the release of OCR 4, a document understanding model designed for enterprise and developer use. This new version brings expanded capabilities, including extraction of structured content with bounding boxes, typed block classification, and inline confidence scores for each region of a document. OCR 4 supports 170 languages across 10 language groups, outperforming previous iterations and other leading systems, particularly with rare and low-resource languages. It is engineered for both high-volume and interactive document workflows, with notable acceleration in processing speed and cost efficiency compared to prior versions and industry competitors. > We ran OCR 4 head-to-head against the field. Independent annotators blindly ranked 600+ real-world documents across 12+ languages, and preferred OCR 4 over every system tested, with win rates averaging 72%. [pic.twitter.com/nGRXtVVQT7](https://t.co/nGRXtVVQT7?ref=testingcatalog.com) > > — Mistral AI (@MistralAI) [June 23, 2026](https://x.com/MistralAI/status/2069420266061475935?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The model is available via API, Mistral Studio, Amazon SageMaker, Microsoft Foundry, and soon on Snowflake Parse Document. For organizations with strict data privacy or residency requirements, OCR 4 can be deployed as a single-container, self-hosted solution. Target customers include enterprises in legal, financial, healthcare, and technical domains that require reliable extraction from complex, multilingual document formats such as PDF, DOC, PPT, and OpenDocument. Mistral’s approach with OCR 4 focuses on delivering precise, localized, and classified document data, enabling downstream use in RAG pipelines, compliance workflows, and enterprise search. Industry engineers have reported substantial reductions in cost and latency when switching to OCR 4, and early users are leveraging the model for structured field extraction, archive digitization, and technical document parsing. [Source](https://mistral.ai/news/ocr-4/?ref=testingcatalog.com) ### ClickUp rolls out Brain² AI with deep workspace context URL: https://www.testingcatalog.com/clickup-rolls-out-brain2-ai-with-deep-workspace-context/ Last updated: 2026-06-23T19:45:04.000Z ClickUp has launched Brain², a full relaunch of ClickUp Brain that transforms the company's built-in AI from an assistant into a context-aware coworker capable of acting across an entire workspace. The release positions Brain as a system that manages the work rather than merely answering questions about it, with every major frontier model available under one subscription and grounded in a team's own data from the first prompt. > The 100x org went viral. Half the internet hated it. The other half was curious. > > One month later: output is up. productivity is spiking. we're approaching a 5:1 agent-to-human ratio. > > And contrary to popular belief, we're doing the OPPOSITE of tokenmaxxing. We're… [pic.twitter.com/MoKZWuJhcR](https://t.co/MoKZWuJhcR?ref=testingcatalog.com) > > — Zeb Evans (@DJ\_CURFEW) [June 23, 2026](https://x.com/DJ%5FCURFEW/status/2069499429292568919?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The central pitch is context. While general-purpose chatbots start each session knowing nothing about a company, Brain² reads tasks, documents, and connected applications in real time and injects that context automatically into whichever model a user selects. Claude, ChatGPT, and Gemini all operate within the same subscription, each with full access to workspace data. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-22-141351.png) Presentation Generated by Brain² Brain selects the model best suited to each step of a task and can switch between them mid-execution. Founder and chief executive Zeb Evans highlighted the gap, stating that general-purpose AI knows nothing about a user's work, whereas Brain² provides any chosen model full access to the workspace so it can act on it, not just describe it. > ClickUp to add artifacts with Brain2 👀 > > \> It will be able to create slides, prototypes, websites, or dashboards. > > \> Brain pulls from workspace context, so the output is built on real project data > Artifacts render inline in the channel and stay fully interactive. > > When Brain is… [pic.twitter.com/ZZhxDFJFPt](https://t.co/ZZhxDFJFPt?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 20, 2026](https://x.com/testingcatalog/status/2068258888382550034?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) In practice, a user can tag Brain in any task, document, or chat and have it read the full thread, pull from connected tools such as Google Drive, GitHub, and Slack via the Model Context Protocol, and then return a finished deliverable rather than a draft answer. > ClickUp's Brain AI will now be able to create agents on its own! > > \> Brain now spots when a task is worth handing off and offers to build a dedicated agent. > > \> It ships preconfigured, with triggers, rules, and scope already in place. > > \> Work keeps moving after Brain finishes, with… [pic.twitter.com/LdLPP3KhaP](https://t.co/LdLPP3KhaP?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 21, 2026](https://x.com/testingcatalog/status/2068617557800530050?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) ClickUp demonstrated the system generating a six-page sales deck from a single tagged lead, and slide decks, dashboards, websites, and working code can each be produced from one prompt. The platform also builds Super Agents, custom AI teammates that run workflows around the clock, and maintains persistent memory of preferences, formatting rules, and team shorthand across sessions instead of resetting each time. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-22-141622.png) ClickUp SuperAgents The retrieval layer is powered by Qatalog, an enterprise search company ClickUp acquired alongside AI coding startup Codegen to build Brain². Qatalog's ActionQuery engine is permission-aware, ensuring Brain² surfaces only what each person is allowed to see, with what the team describes as zero index lag across more than one hundred integrations. Brain also references the sources behind each response so its work can be verified, and it includes a system prompt designed to question decisions rather than simply agree with them. ClickUp lists ISO 42001 certification and zero data retention for model training as part of the package. SPONSORED Start testing Brain² on Clickup! [Learn more ](https://clickup.com/brain?ref=testingcatalog.com) ClickUp has been advancing into AI over the past two years, raising more than half a billion dollars, and has expressed its intention to go public. Brain² extends that work-management platform into a model-agnostic AI layer that competes with offerings from Notion, Slack, and Asana, and is available on desktop and mobile for teams already working in ClickUp. ### Latitude launches open-source platform to monitor AI agents URL: https://www.testingcatalog.com/latitude-launches-open-source-platform-to-monitor-ai-agents/ Last updated: 2026-06-23T15:55:25.000Z Latitude has released an open-source platform for monitoring AI agents in production, built to show what an agent is doing once it meets real users, catch where it breaks down, and route the fix back to the editor where the code already lives. At the base is a discovery layer that gathers thousands of live conversations and clusters them into one picture of what people ask for and where they hesitate, escalate, or drop off. Usage can be broken down by who is behind it, from power users to one-time visitors to the accounts hitting failures most often, and individual sessions can be inspected alongside their cost, latency, and problems. A semantic search lets a team type a question in plain language, such as where users mention a feature the agent does not offer, and the matching conversations come back directly. The second layer turns scattered breakages into something a team can act on. When an agent keeps failing the same way, Latitude collapses those moments into a single signal that names the problem, counts how often it occurs, and attaches the likely reason. Signals come from automatic flaggers, annotations, or manual creation, and an evaluation is generated for each one. A saved search can be promoted into a monitor that runs against every new conversation, so a pattern reaches the team before it reaches more users. ![Latitude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/TeYtCvXdVcvagk8JLijif3pvZFU.webp) The third layer closes the loop inside the developer workflow. An MCP server delivers projects, traces, signals, searches, and datasets directly to a coding agent, so the work happens where engineers already operate rather than in a separate console. Production conversations can be turned into datasets and reused as test sets, allowing a team to confirm that a fix holds before shipping. The platform also connects to Claude Code to track token spend per task, surfacing where the budget goes. Latitude is distributed under an MIT license and can run on a team's own infrastructure, with a free tier and full source access for anyone who wants to read or modify it. The framing treats an agent as the richest record a company holds about its own product, and the platform as the way to read that record back rather than leaving it unused. SPONSORED Star the Latitude repo! [Learn more ](https://github.com/latitude-dev/latitude-llm?ref=testingcatalog.com) Latitude is the team behind the open-source project of the same name, maintained under the latitude-dev organization on GitHub. The release reflects a move toward treating agent monitoring as a loop in which the system reports what went wrong and points to the fix, rather than a dashboard to be watched, and it arrives as agent reliability becomes a defining concern for teams shipping to real users at scale. ### OpenAI prepares bidirectional voice mode for rollout on ChatGPT URL: https://www.testingcatalog.com/openai-prepares-bidirectional-voice-mode-for-rollout-on-chatgpt/ Last updated: 2026-07-17T21:53:19.000Z OpenAI looks set to hand ChatGPT's voice mode its biggest upgrade in months, with a next-generation audio model surfacing as [Bidi 1](https://www.testingcatalog.com/openai-prepares-major-chatgpt-voice-upgrade-with-gpt-bidi-1/), shorthand for the bidirectional design that lets the assistant speak, hear, and listen at once. References to it began appearing in the ChatGPT web interface ahead of a possible release this week, and it has already begun reaching a subset of users in the app. > BREAKING 🔥: First tests of "Bidi 1", an upcoming bidirectional voice model from OpenAI. This upgrade will arrive in ChatGPT and, potentially, in Codex soon as well. > > \> Bidi 1 can speak over while you are talking and keep listening. > \> Bidi 1 can switch between tasks back and… [https://t.co/BwWhCKx3G0](https://t.co/BwWhCKx3G0?ref=testingcatalog.com) [pic.twitter.com/Fawc74kBym](https://t.co/Fawc74kBym?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 23, 2026](https://x.com/testingcatalog/status/2069331697615749530?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) In our early testing, the gap from today's advanced voice mode is plain. Bidi 1 sits in the model selector under settings, beside the standard and advanced options, and turns the voice bubble yellow once picked. It offers small, natural acknowledgments — an "okay" or a brief nod — when you pause or slow down, without cutting across you. It also switches tasks on the fly: ask it to count to ten, interrupt to reverse the count, and it adjusts immediately. > OPENAI 🔥: An upcoming Bidi 1 voice model will be able to translate in real-time! > > This will unlock a huge pile of use cases to be built on top of when it lands on the APIs. [pic.twitter.com/95sRnSzJfs](https://t.co/95sRnSzJfs?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 23, 2026](https://x.com/testingcatalog/status/2069351216648204757?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) More usefully, it holds the thread of a whole conversation rather than dropping earlier context, the weak point that has long dogged the current voice stack, and it no longer jumps in during longer pauses. ![ChatGPT](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/ChatGPT-06-23-2026_02_41_AM.jpg) Creative behavior carries over from the first advanced voice rollout, singing and beatboxing included, though copyright handling is tighter; it declines popular songs outright while still attempting an original piece in a chosen artist's style. The move reads as [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) closing the distance between its capable text models and an older voice layer, treating conversation as a core route into ChatGPT. The company has not formally announced it. A gradual, opt-in release across web and mobile looks likely, with the European Economic Area possibly waiting longer (not confirmed). Codex appears set for its own voice upgrade in the weeks after this launch, separate from it, and API access may follow later still (timeline is not confirmed). ### Anthropic prepares Cowork support for mobile apps URL: https://www.testingcatalog.com/anthropic-prepares-cowork-support-for-mobile-apps/ Last updated: 2026-06-23T00:13:12.000Z Anthropic appears close to extending Cowork, its agentic system for knowledge work, beyond the desktop. A build of the iOS app carries a Cowork entry, gated behind a feature flag, that would surface in the side navigation. Tapping it reveals copy about scheduling Cowork tasks from a phone and picking up results across mobile, web, or desktop, plus a tab that gathers every scheduled action in one place. > Users will be able to trigger Cowork tasks on mobile and view scheduled tasks in the app. > > Since we got announcements being added to the app, we may see it already dropping this week. [pic.twitter.com/fDLNWTHn5x](https://t.co/fDLNWTHn5x?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 22, 2026](https://x.com/testingcatalog/status/2069095762793787563?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) That framing points to a shift rather than a debut. Cowork already reached phones in March through Dispatch, which lets someone message a desktop session remotely, but only while the computer stays awake, since the work runs locally. The new wording suggests moving execution to the cloud and the web, lifting Dispatch's central constraint so tasks can run without a machine left on. The copy is already written, hinting at a release as soon as this week, though nothing is operational yet. A second finding concerns voice. New consent text in the app sources accompanies a [model selector for voice mode](https://www.testingcatalog.com/claude-code-managed-agents-and-model-selector-for-voice-mode/), letting users choose the model behind the spoken experience. [Claude](https://www.testingcatalog.com/tag/claude/) Voice has leaned on Haiku 4.5 for a while, so the selector reads as groundwork for an underlying model refresh, on top of the multilingual rollout already underway across accounts. ### Google tests literature review matrix tool for NotebookLM URL: https://www.testingcatalog.com/google-tests-literature-review-matrix-tool-for-notebooklm/ Last updated: 2026-06-22T22:40:31.000Z Google is preparing a new artifact type for NotebookLM that would generate a literature review matrix from a user's uploaded sources. Labeled "Lit Review" in its current pre-release form, a name that may shift before launch, it appears built to turn a stack of documents into a structured comparison grid rather than the prose reports the tool already produces. It is not yet available to anyone, and no timeline has surfaced. The matrix format signals who Google is courting: students, academics, and anyone working closely with large bodies of text. A grid lining up themes, arguments, or methods across sources is a staple of formal research, and placing it inside NotebookLM's Studio panel would give that audience a one-click way to map a field. The structure carries beyond academia. Feed in a book or a full series, and it could, in principle, chart characters, plots, or motifs against one another. "A Song of Ice and Fire" is the kind of sprawling, multi-volume corpus that would put such a matrix through its paces. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/NotebookLM-06-23-2026_12_39_AM.jpg) That reading-focused angle aligns with a second thread of development: a Google Play Books bridge and a dedicated textbooks section among [NotebookLM's](https://www.testingcatalog.com/tag/notebooklm/) source options. The app already accepts EPUB files, and a [Play Books](https://www.testingcatalog.com/notebooklm-tests-mind-map-controls-and-play-books-sources/) link would let rights-protected reading material flow into a grounded workflow without manual copying. The moves fit Google's recent direction, which moved NotebookLM to Gemini 3.5 this month alongside agentic research, code-running notebooks, and broader downloadable outputs, tightening the connective tissue between Google's reading, research, and Workspace surfaces. Caveats apply. The tool's source-grounded summaries have historically slipped on citation accuracy, and a matrix is only as reliable as the mapping behind it. For now, it sits in development, its final form and arrival date both open. ### OpenAI launches new security tools and updates GPT-5.5-Cyber URL: https://www.testingcatalog.com/openai-launches-new-security-tools-and-updates-gpt-5-5-cyber/ Last updated: 2026-06-22T21:42:01.000Z OpenAI is advancing Daybreak beyond vulnerability discovery into patch automation, launching an updated Codex Security plugin, the full GPT-5.5-Cyber model in limited release, a Daybreak Cyber Partner Program, and Patch the Planet, an open-source security initiative built with Trail of Bits, HackerOne, Calif, researchers, and project maintainers. The core shift is from finding bugs to landing fixes. Codex Security is integrated into Codex workflows and can scan an entire codebase, a selected folder, or a specific change. It can review recent commits, produce reports with severity, affected code locations, validation evidence, and remediation guidance, trace attack paths, build threat models, validate findings, generate patches, and export results into vulnerability management systems through formats such as SARIF and CodeQL queries. Since its research preview in March, OpenAI reports that Codex Security scanned more than 30 million commits across over 30,000 codebases, with human reviewers marking more than 70,000 findings as fixed and over 500,000 findings automatically detected as fixed. > We want to help all companies be secure, working with the USG and the security ecosystem. > > \*The full version of GPT-5.5-Cyber is here; state of the art performance on CyberGym. > > \*Patch The Planet and Codex Security will help solve security problems instead of just finding them. [pic.twitter.com/otyCFHJR4d](https://t.co/otyCFHJR4d?ref=testingcatalog.com) > > — Sam Altman (@sama) [June 22, 2026](https://x.com/sama/status/2069121360744550796?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) GPT-5.5-Cyber is the more controlled but more capable part of this release. OpenAI states that the model is intended for verified defenders working on authorized cybersecurity tasks, not general access. It is designed for deeper codebase analysis, reachability checks, vulnerability validation, patch development, testing, and evidence preparation. On CyberGym, GPT-5.5-Cyber reached 85.6 percent compared with 81.8 percent for GPT-5.5\. It also scored 39.5 percent on ExploitGym versus 25.95 percent for GPT-5.5, and 69.8 percent on SEC-bench Pro versus 63.1 percent. Patch the Planet brings this capability into open-source software. More than 30 projects have committed to participate, with initial names including cURL, Go, Python, Sigstore, pyca/cryptography, NATS Server, aiohttp, freenginx, and Python.org. Participating projects receive ChatGPT Pro, conditional access to Codex Security, and API credits for maintainer automation and release workflows. Trail of Bits engineers are working directly with maintainers to validate issues, remove duplicates, reassess severity, write patches, support tests, and coordinate disclosure before maintainers see the final work. OpenAI is also promoting Daybreak through a partner model rather than direct broad model access. The aim is to embed GPT-5.5 with Trusted Access for Cyber into existing security products and services, while keeping access governed through partner systems. The company is positioning Daybreak as a defensive cyber stack for the AI era: frontier models, Codex workflows, controlled access, expert review, and security ecosystem integrations. The release is significant because OpenAI is no longer presenting AI cybersecurity only as a model capability or evaluation result. It is transforming it into an operational pipeline for scanning, validating, fixing, and reviewing software vulnerabilities across enterprise, government, and open-source environments. [Source](https://openai.com/index/daybreak-securing-the-world/?ref=testingcatalog.com) ### Sakana AI releases Fugu Ultra system to rival top AI labs URL: https://www.testingcatalog.com/sakana-ai-releases-fugu-ultra-system-to-rival-top-ai-labs/ Last updated: 2026-06-22T21:41:44.000Z Sakana AI has announced the launch of Fugu Ultra, a frontier-level orchestration model designed to rival top-tier AI models such as Anthropic's Fable 5 and Mythos Preview. This new release is available to the general public worldwide through a single, OpenAI-compatible API, making it accessible for both individual and enterprise customers. Fugu Ultra is tailored for users with demanding, multi-step tasks across engineering, science, research, cybersecurity, and data analysis, addressing the need for resilient AI infrastructure that avoids single-vendor risks or export control restrictions. > Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. > > Our ‘Fugu Ultra’ model matches the performance of Fable and Mythos, delivering frontier capability without the risk of export controls. > > Try it: [https://t.co/aDEFyySWlS](https://t.co/aDEFyySWlS?ref=testingcatalog.com) 🐡 [pic.twitter.com/43wzMAhyzT](https://t.co/43wzMAhyzT?ref=testingcatalog.com) > > — Sakana AI (@SakanaAILabs) [June 22, 2026](https://x.com/SakanaAILabs/status/2068861630327443966?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Fugu Ultra distinguishes itself by utilizing a multi-agent orchestration system: it is a language model that can autonomously delegate subtasks to a pool of expert LLMs, including instances of itself, choosing and combining their outputs for the highest quality results. Benchmark comparisons show that Fugu Ultra matches or surpasses leading competitors on coding, reasoning, and scientific tasks. Early user feedback highlights its strengths in sustained persona consistency and thoroughness in long, complex workflows, such as code review and security assessment, where it surfaces more issues and maintains focus better than other models. ![Fugo](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/benchmark-fugu-grid.png) Sakana AI positions this release as a step toward AI sovereignty, offering ongoing improvements as new models are added to the agent pool. The company, established with a mission to build collaborative AI ecosystems rather than monolithic models, emphasizes Fugu Ultra’s role in providing organizations with greater operational and geopolitical security. [Source](https://sakana.ai/fugu-release/?ref=testingcatalog.com) ### Perplexity releases Brain Memory System for Perplexity Computer URL: https://www.testingcatalog.com/perplexity-releases-brain-memory-system-for-perplexity-computer/ Last updated: 2026-06-19T21:39:41.000Z Perplexity released a new memory system, referred to internally as Brain, that would sit beneath several of its core products rather than living inside any single one. Signals point to Brain powering Perplexity Search, the Perplexity Computer agent, and possibly the Comet browser, positioning memory as a shared connective layer instead of a per-product add-on. > With Brain, Computer starts each task with full context of your projects, decisions, and sources instead of from scratch. > > On tasks that require past context, Brain improves answer correctness by 25%, recall by 16%, and runs 13% cheaper per task. [pic.twitter.com/xlNHTmA1My](https://t.co/xlNHTmA1My?ref=testingcatalog.com) > > — Perplexity (@perplexity\_ai) [June 18, 2026](https://x.com/perplexity%5Fai/status/2067642159793406112?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) What sets the approach apart, based on what has surfaced so far, is transparency. Where most assistant memory operates as an opaque store, Brain looks set to expose the full knowledge base to the user in three forms: 1. Topics sorted into categories 2. The underlying context held behind each topic 3. A navigable 3D map of the connections between them that a person can hover over and explore ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Perplexity-Computer-Memory-06-19-2026_11_24_PM.jpg) The clustering follows the now-familiar second-brain pattern, closer to an Obsidian-style graph than a flat list, where the model gathers information and groups it by subject. Because context is organized by topic and retrieved only when a task calls for it, the amount pulled into any given request appears lower, which would help both speed and recall quality. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Perplexity-Computer-Memory-06-19-2026_11_25_PM.jpg) The company has spent the year moving from answering questions toward doing work, with the Computer agent arriving in late February alongside a memory upgrade that sharpened recall accuracy. Brain would extend that arc, turning memory into the substrate that lets search, agents, and browsing draw on the same accumulated context. ![Perplexity](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Perplexity-Computer-Memory-06-19-2026_11_26_PM.jpg) [Perplexity](https://www.testingcatalog.com/tag/perplexity/) Brain also lands in a crowded field: always-on agents such as Hermes and OpenClaw lean on their own knowledge bases, and rivals from Claude to Notion are building comparable persistent stores. The feature remains hidden for now, but the recent pace of polishing around it suggests a wider rollout may not be far off. [Official announcement](https://www.perplexity.ai/hub/blog/self-improving-memory-for-agents?ref=testingcatalog.com) ### Anthropic launches live Artifacts for Claude Code URL: https://www.testingcatalog.com/anthropic-launches-live-artifacts-for-claude-code/ Last updated: 2026-06-19T21:38:17.000Z Anthropic has announced the launch of Artifacts in Claude Code, a new capability allowing users to generate live, shareable visual pages directly from their coding sessions. This feature is available in beta for Claude Team and Enterprise organizations, accessible via the Claude Code CLI and desktop app, with artifacts viewable in any browser. Artifacts transform the output of coding sessions, ranging from incident investigations to code refactoring, into interactive web pages such as PR walkthroughs, dashboards, and release checklists that update in real time as the session progresses. Each page is built from the full session context, including codebase details, connectors, and conversation history, eliminating the need for manual data integration or additional infrastructure. Version control ensures each artifact retains a history, and a gallery feature enables easy management of all generated artifacts. > New in Claude Code: Artifacts. > > Interactive pages built from your session, like a PR walkthrough or a living project dashboard, shared with your team at a private link. > > Available in beta on Team and Enterprise plans. [pic.twitter.com/0NX9gNCaAs](https://t.co/0NX9gNCaAs?ref=testingcatalog.com) > > — Claude (@claudeai) [June 18, 2026](https://x.com/claudeai/status/2067671912038240487?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Anthropic, the company behind Claude Code, is known for its focus on AI-powered developer tools designed to streamline workflows and support teams working on complex software projects. With Artifacts, Anthropic targets engineering teams, SREs, architects, managers, and other technical roles seeking better ways to document and share ongoing work. Access controls ensure all artifacts remain private to the organization, with admins able to manage permissions and retention through org-level settings. Early feedback from internal testing highlights the benefits for debugging and collaborative incident response, as team members can instantly view up-to-date information without waiting for manual updates. [Source](https://claude.com/blog/artifacts-in-claude-code?ref=testingcatalog.com) ### Anthropic launches managed connector access with Okta URL: https://www.testingcatalog.com/anthropic-launches-managed-connector-access-with-okta/ Last updated: 2026-06-19T21:33:29.000Z Anthropic has introduced enterprise-managed authorization for MCP connectors, enabling organization-wide provisioning through identity providers, beginning with Okta. This update allows administrators to centrally configure connector access, eliminating the need for individual user authorization. Upon first login, users automatically receive access to the designated connectors based on their IdP group assignments, ensuring connectors are available from the outset across Claude chat, Claude Code, and Cowork. > We've added support for the Enterprise-Managed Auth extension to MCP. > > Admins can centrally authorize MCP connectors for their organization, so all the tools and data users need are connected on their first login. [pic.twitter.com/NJl57QbxD0](https://t.co/NJl57QbxD0?ref=testingcatalog.com) > > — ClaudeDevs (@ClaudeDevs) [June 18, 2026](https://x.com/ClaudeDevs/status/2067655887662272723?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The feature is built on an open standard as part of the Enterprise-Managed Authorization extension to the Model Context Protocol, making it compatible with both existing and custom connectors. It is designed for enterprise customers who manage large teams and require streamlined, secure access management integrated with existing security workflows. Supported MCP providers at launch include: 1. Asana 2. Atlassian 3. Canva 4. Figma 5. Granola 6. Linear 7. Supabase Slack integration is planned. Organizations such as Hubspot, Ramp, and Webflow are already deploying this system to their teams. Admins benefit from unified access control, as connector permissions and revocation are administered through the trusted identity provider, minimizing the risk of lingering access. The approach also enforces separation between work and personal accounts, supporting compliance and security requirements. This positions [Anthropic’s](https://www.testingcatalog.com/tag/claude/) solution as a comprehensive response to the need for scalable, secure connector management in enterprise environments. [Source](https://claude.com/blog/enterprise-managed-auth?ref=testingcatalog.com) ### OpenAI prepares real-time voice mode for Pets in Codex URL: https://www.testingcatalog.com/openai-prepares-real-time-voice-mode-for-pets-in-codex/ Last updated: 2026-06-19T00:54:15.000Z OpenAI looks set to bring real-time voice to Codex, and the way it is arriving says as much about branding as about capability. A voice control that once did nothing now summons the animated pet the coding app gained earlier this year, a sign that the session handling has come online ahead of a public switch. The timing lines up with separate signals that a next-generation voice model, tentatively tagged [GPT-Bidi-1](https://www.testingcatalog.com/openai-prepares-major-chatgpt-voice-upgrade-with-gpt-bidi-1/), could reach ChatGPT's voice mode next week, built around a bidirectional design that listens and speaks at once rather than taking strict turns. 0:00 /0:33 1× Realtime voice button triggering Codex Pet Hoots Inside the app, a real-time voice section now includes controls for developers who want to keep the spoken channel open while code runs. One assigns a hotkey and a wake word, with sessions starting on the phrase "Hey Chat." A single-tone option pins new sessions to one durable orchestrator thread rather than spawning fresh ones, letting a developer return to the same continuous context. ![Dedicated settings tab for Pets in Codex](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-18-at-23.26.20.webp) Dedicated settings tab for Pets in Codex A third setting governs the avatar overlay, an Orb or the [Pet](https://www.testingcatalog.com/openai-adds-animated-pets-and-config-imports-to-codex/), though the Orb has yet to render since the Pet keeps appearing. A Library entry has also surfaced in the sidebar, not yet openable but closely modeled on the one ChatGPT already uses. ![Library on Codex](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-18-at-23.42.16.webp) Library on Codex That borrowing is the real story. The orb, the library, and a wake word that says "Chat" rather than "Codex" suggest OpenAI treats its coding agent and consumer assistant as one converging surface, not two products. For a company betting speech becomes the main route to its models, a single surface is the point. Whether the Codex side flips the same week as the ChatGPT model or trails behind, which is plausible, the direction is set. Anthropic is working on the same ground. [Multilingual support](https://www.testingcatalog.com/anthropic-plans-expanding-claude-voice-mode-to-more-languages/) has begun appearing in Claude's voice mode without a formal announcement, alongside a push-to-talk option, and it reads as a first layer rather than the full plan, with the [underlying model and system](https://www.testingcatalog.com/claude-code-managed-agents-and-model-selector-for-voice-mode/) likely due for an upgrade, too. The quiet rollout fits a habit of holding the louder reveal until a rival forces the moment. ### OpenAI prepares GPT-5.6 models for the upcoming release URL: https://www.testingcatalog.com/openai-prepares-gpt-5-6-models-for-the-upcoming-release/ Last updated: 2026-06-18T20:23:56.000Z OpenAI looks set to widen its GPT-5 line next week, with the GPT-5.6 family poised for a Tuesday arrival. The plan appears to span the standard model, potentially alongside Mini and Pro variants, landing together rather than trickling out, though the company's recent habit of staggering ChatGPT, API, and Pro releases across several days leaves room for doubt on that point. > OPENAI 🔥: GPT-5.6 and GPT-5.6-Pro models may potentially arrive as soon as next week. > > Really soon 👀 [https://t.co/MspMlB3SMR](https://t.co/MspMlB3SMR?ref=testingcatalog.com) [pic.twitter.com/XjIZ5cA6lR](https://t.co/XjIZ5cA6lR?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 18, 2026](https://x.com/testingcatalog/status/2067652120673739048?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Early traces of a GPT-5.6 Pro build have surfaced for some Pro subscribers, and the first outputs we managed to pull looked strong. > GPT-5.6 Pro a généré ce SVG de Windows 11 🔥 > > Franchement, je le trouve meilleur que Mythos sur ce prompt. > > Le problème, il ajoute des éléments inutiles (pop-ups, beaucoup de textes, etc.). [https://t.co/oOTD6WNVG7](https://t.co/oOTD6WNVG7?ref=testingcatalog.com) [pic.twitter.com/N5L9pZalzY](https://t.co/N5L9pZalzY?ref=testingcatalog.com) > > — Mirochill (@mirochill) [June 18, 2026](https://x.com/mirochill/status/2067686414636941540?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > 🚨 GPT 5.6 Pro first output on the same prompt > > we are getting started > \> frontend/ webdev is not solved or improved yet > \> but understanding increased a lot > \> it started to take 20-40 mins again like it used to do before 5.5 pro [https://t.co/zcLehTbe5c](https://t.co/zcLehTbe5c?ref=testingcatalog.com) [pic.twitter.com/C7u6ZRUfjT](https://t.co/C7u6ZRUfjT?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [June 18, 2026](https://x.com/chetaslua/status/2067642369403740220?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > GPT 5.6 Pro continues to mogs Fable in 3d test > > working on games one shot too , for comment check quote post and also share your prompts and fable output i will tag you with 5.6 pro results > > credit : [@mirochill](https://x.com/mirochill?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [https://t.co/kSMkWpOnyZ](https://t.co/kSMkWpOnyZ?ref=testingcatalog.com) [pic.twitter.com/6lsp3pUC6M](https://t.co/6lsp3pUC6M?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [June 18, 2026](https://x.com/chetaslua/status/2067654353318806008?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The competitive angle is hard to miss. Chatter places GPT-5.6 against Anthropic's top tier, with developers claiming it edges out the Mythos line on agentic coding work. Rumored gains center on a context window pushed toward 1.5 million tokens, up from the million GPT-5.5 shipped with in April, plus sharper long-horizon coding and quicker Codex response times. Pricing is the quieter story: OpenAI already undercuts Anthropic by roughly half on tokens, and reports point to deeper cuts aimed at a price war. Timing sharpens the stakes. Anthropic's Claude Fable 5, a Mythos-tier model, is [subject to a US regulatory action](https://www.testingcatalog.com/anthropic-suspends-fable-5-and-mythos-5-after-export-order/) that has left its availability uncertain. A window has opened, and OpenAI looks intent on using it. Whether Washington's posture shifts again is the open variable, but the launch calendar reads as set. SPONSORED [![CTA Image](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Discord-06-18-2026_10_23_PM.jpg)](https://discord.com/invite/devmode?ref=testingcatalog.com) Check DevMode Discord for more 👀 [Join ](https://discord.com/invite/devmode?ref=testingcatalog.com) The family is not the only thing in motion. A next-generation voice model, working name [GPT-Bidi-1](https://www.testingcatalog.com/openai-prepares-major-chatgpt-voice-upgrade-with-gpt-bidi-1/), is also expected to land, bringing a bidirectional audio system built to listen and speak at once, ride over interruptions, and adjust mid-sentence. Inside ChatGPT, it would sit beside today's Advanced Voice Mode as a separate option, carrying High, Medium, and Instant tiers that mirror the text side. A draggable voice bubble spotted in the interface appears to be an early piece of that redesign. ### Exclusive: Microsoft evaluates different open models for Cowork URL: https://www.testingcatalog.com/exclusive-microsoft-evaluates-different-open-models-for-cowork/ Last updated: 2026-06-18T14:43:51.000Z TestingCatalog has learned that **Microsoft teams behind Copilot Cowork are evaluating a broader set of open and open-weight models beyond DeepSeek** as potential underlying options for the agentic work app. [Axios recently reported](https://www.axios.com/2026/06/16/microsoft-copilot-cowork-tokenmaxxing-cowork?ref=testingcatalog.com) that Microsoft is considering a Microsoft-hosted version of DeepSeek as a cheaper model option for Copilot Cowork, but the evaluation appears to extend beyond a single model family. > New [@axios](https://x.com/axios?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com): Microsoft eyes DeepSeek for Copilot Cowork as it also joins the shift to usage based pricing. Says final decision TK but it has already fine-tuned a model that it could use. > > — Ina Fried (@inafried) [June 16, 2026](https://x.com/inafried/status/2066921423378190763?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The key architectural detail is the separation between the model layer and the harness. Copilot Cowork is being built so the orchestration system can remain stable while the underlying models are swapped depending on task type, cost, latency, and required capability. That would allow Microsoft to route some work to frontier APIs, while other parts could be handled by cheaper self-hosted models on Azure. Over time, some lighter tasks may also become candidates for local execution as smaller models mature. This direction would make sense for [Copilot Cowork](https://www.testingcatalog.com/exclusive-new-screenshots-of-upcoming-copilot-super-app/), which is now [generally available](https://www.testingcatalog.com/microsoft-launches-copilot-cowork-globally-for-microsoft-365/) for Microsoft 365 Copilot customers and is aimed at enterprise users who expect reliability, governance, and data controls. For customers, the main benefit would be more flexible pricing and model choice. For Microsoft, it offers a way to reduce dependence on any single external model provider while controlling compute costs for long-running agentic workflows. At the same time, this creates an internal competition. Microsoft has been presenting its own [MAI models](https://www.testingcatalog.com/microsoft-build-2026-recap-from-windows-to-copilot-all-ai/) as part of a broader push to become a top-tier AI lab next to OpenAI, Anthropic, and Google. If Chinese and open-weight models continue to perform strongly enough to power parts of Copilot Cowork, Microsoft’s internal model teams will face a clearer benchmark: they need to compete with fast-moving open-model providers. **According to sources familiar with the evaluation, these open-model developments have not reached production in Copilot Cowork yet**. For now, they remain under active testing. The practical question is not whether Microsoft can plug in another model, but which model can meet enterprise expectations once cost, compliance, safety, and task quality are measured together. ### Zeta Labs brings AI employee Viktor to Microsoft Teams URL: https://www.testingcatalog.com/zeta-labs-brings-ai-employee-viktor-to-microsoft-teams/ Last updated: 2026-06-18T14:04:53.000Z Zeta Labs has brought Viktor, its AI employee, into Microsoft Teams, extending a product that already runs inside more than 20,000 Slack workspaces. Until now, Viktor operated only on Slack, where it sits in channels and works alongside a team rather than waiting in a separate window for a prompt. The Microsoft Teams release opens the same capability to organizations that run their daily communication on Microsoft's platform, and the company frames the move around one idea: Teams is where people talk about work, and Viktor turns those conversations into completed tasks. Viktor is positioned as an AI employee rather than an assistant or a chatbot. It joins a channel that maintains a persistent record of what the team has already done, maintains context for what the group is working toward, and acts before every instruction is spelled out. Rather than returning answers on request, it writes and runs its own code to produce finished output, from reports and dashboards to web apps. One license covers everyone in a channel, so the same instance works for the whole team rather than being tied to a single person. > Excited to announce Viktor in Microsoft Teams. > > This week we crossed $20M in annualized revenue run rate. > > In Slack. One app, no sales team, no rollout. > > Now Viktor goes where the rest of the working world actually is. > > 320 million people work in Microsoft Teams. The biggest org… [pic.twitter.com/iaAJzxChNJ](https://t.co/iaAJzxChNJ?ref=testingcatalog.com) > > — Fryd Wiatrowski (@frydwia) [June 18, 2026](https://x.com/frydwia/status/2067601589011914989?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The footprint centers on connectivity and memory. Viktor reads from and writes to more than 3,000 tools, including Salesforce, HubSpot, Stripe, GitHub, Google Ads, Notion, and Linear, which lets it move information and act across the systems a team already runs. Because it retains memory across sessions, it does not reset between conversations or need a team to re-explain its company each morning. Zeta Labs reports SOC 2 Type 1 certification and says Viktor is approved by Microsoft for Teams, two points aimed at organizations with security and procurement review before they adopt an autonomous agent. The intended audience covers operations teams, agencies, and growing businesses that want recurring work handled without adding headcount, the same profile Viktor built on Slack. Early customers describe it as quickly adopted across a team and able to automate manual processes that previously did not scale, and the company is careful to frame Viktor as additive to a team rather than a replacement for it. SPONSORED Start testing Viktor [Learn more ](https://viktor.com/?ref=testingcatalog.com) Zeta Labs was founded by Fryderyk Wiatrowski and Peter Albert, former Meta engineers, and builds on the autonomous agent infrastructure behind its earlier product, Jace AI. The company reports a revenue run rate above $15 million, and in May 2026, it raised $75 million in Series A funding led by Accel, bringing total funding to $77.9 million. The Microsoft Teams release is the company's second major platform after Slack, putting one AI employee inside the two environments where most workplace conversations happen. New accounts receive $100 in credits at signup, and Viktor is now available on the company website. ### Mistral AI to get Code and Apps features on Vibe URL: https://www.testingcatalog.com/mistral-ai-to-get-code-and-apps-features-on-vibe/ Last updated: 2026-06-18T11:38:10.000Z Mistral AI looks set to reshape the web version of Vibe (Le Chat) with two additions spotted in development, both pointing beyond pure conversation. The first is a **Code** section, positioned as a peer to the existing Chat and Work areas. The company's coding agents have so far centered on its command-line tool, and the web build appears to bring that into the browser, likely mirroring the CLI. Whether it also foreshadows a desktop client remains unclear, though its placement suggests that coding is becoming a first-class surface rather than a side feature. Developers who would rather skip terminal setup are the obvious early audience. The second, still flagged as a work in progress, reads as [Mistral's](https://www.testingcatalog.com/tag/mistral/) take on advanced artifacts: an **Apps** area. Users would be able to build, host, and share apps that pull data through connectors or run multi-step workflows, moving Le Chat from somewhere to ask questions toward somewhere to ship tools. That tracks with Mistral's recent connector directory and its Workflows engine, and would place it alongside Anthropic's shareable artifacts and OpenAI's in-chat apps. Real limits are hard to gauge while the feature is unfinished. > This model and upcoming ones will be open-weight. We believe this is critical for our customer confidence and for the research and developer communities. You cannot own, inspect, audit, or improve a system you are only permitted to reach through someone else's interface,… > > — Arthur Mensch (@arthurmensch) [June 16, 2026](https://x.com/arthurmensch/status/2066913359409090967?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The timing is notable, as Arthur Mensch has confirmed a new model arriving this summer, described as the start of a fresh family that is large yet sparse, wording that points to a mixture-of-experts design. He said it will ship as open weights, with early access opening in July for partners across research, government, and industry. Together, the browser features and the summer model sketch a wider shift: Mistral moving from a model lab toward a full product platform, leaning on its European, openly licensed footing as the line against larger US rivals. It makes for a crowded summer, and the release order will say a lot about where the company sees its leverage. ### ICYMI: ZAI launches GLM-5.2 open model with 1M context URL: https://www.testingcatalog.com/icymi-zai-launches-glm-5-2-open-model-with-1m-context/ Last updated: 2026-06-18T00:47:53.000Z Z.ai has released GLM-5.2, a new flagship text model built for long-horizon coding agents, project-scale software work, automated research, debugging, refactoring, mobile development, and code-driven video generation. The model offers a 1M-token context window and up to 128K output tokens, positioning it for full-repository tasks where it needs to retain architecture, API contracts, file boundaries, prior decisions, and engineering rules across long sessions. > Introducing GLM-5.2: Frontier Intelligence, Open Weights > > \- Significant improvements in coding and agentic tasks > \- Strong long-horizon capabilities with a 1M context window > \- Two levels of reasoning effort: GLM-5.2 (max) pushes the limits, while GLM-5.2 (high) strikes a strong… [pic.twitter.com/SjGPSVhePJ](https://t.co/SjGPSVhePJ?ref=testingcatalog.com) > > — Z.ai (@Zai\_org) [June 16, 2026](https://x.com/Zai%5Forg/status/2066938937344495629?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) GLM-5.2 is available to GLM Coding Plan users across Lite, Pro, and Max tiers, with switching support inside coding agents such as Claude Code, OpenClaw, and Cline through custom model configuration. Developers can enable the 1M-token version with the `glm-5.2[1m]` model name, map Claude Code effort modes to GLM-5.2’s high or max reasoning levels, and use Z.ai’s OpenAI-compatible API endpoint for integrations. The core upgrade is not just context length. Z.ai says GLM-5.2 was trained for long-horizon coding agent scenarios, including large-scale implementation, automated research, performance optimization, and complex debugging. The company claims GLM-5.2 trails Claude Opus 4.8 by 1 percentage point on FrontierSWE, beats GPT-5.5 and Opus 4.7 on multiple long-horizon benchmarks, and scores 81.0 on Terminal-Bench 2.1 versus 62.0 for GLM-5.1\. On SWE-bench Pro, it scores 62.1, compared to 58.4 for GLM-5.1. ![ZAI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/20260617-012836.webp) The model also adds architectural changes for long-context inference. Z.ai says GLM-5.2 uses IndexShare, a sparse-attention method that reuses the same indexer across every 4 sparse-attention layers, reducing per-token FLOPs by 2.9x at a 1M context length. The company also says it updated the model’s MTP layer for speculative decoding, increasing acceptance length by up to 20%. Download listings show GLM-5.2 and GLM-5.2-FP8 builds with 744B total parameters and 40B active parameters on Hugging Face and ModelScope. Early developer feedback cited by Z.ai centers on project-level context capacity, steadier long-running execution, stricter adherence to production engineering constraints, and stronger client-side and mobile workflows, including ADB, logcat, screenshots, runtime logs, Mini Program migration, and real-device debugging loops. That makes GLM-5.2 less of a chatbot release and more of a direct play for coding-agent infrastructure. Z.ai, formerly ZhipuAI, is the company behind the GLM model family. It was founded in 2019 from Tsinghua University's technological work and has since released GLM, GLM-130B, ChatGLM, GLM-4, GLM-4-Voice, AutoGLM, and other agent and model products. The company says ChatGLM-6B has surpassed 20 million global downloads, giving GLM-5.2 a clear role in its push from open-model research to developer-facing AI coding infrastructure. [Source](https://z.ai/blog/glm-5.2?ref=testingcatalog.com) ### Microsoft launches Copilot Cowork globally for Microsoft 365 URL: https://www.testingcatalog.com/microsoft-launches-copilot-cowork-globally-for-microsoft-365/ Last updated: 2026-06-18T00:27:22.000Z Microsoft has announced the worldwide general availability of [Copilot Cowork](https://www.testingcatalog.com/exclusive-new-screenshots-of-upcoming-copilot-super-app/), following a three-month preview with strong adoption from major enterprises, including over half of the Fortune 500\. This release targets Microsoft 365 Copilot customers and is accessible globally to organizations with a Copilot User Subscription License. Copilot Cowork is aimed at corporate knowledge workers, managers, customer-facing staff, and technical teams seeking to automate complex, multi-step tasks across documents, data, and workflows. > Copilot Cowork is now generally available! > > Over the last few months of preview in Frontier, we’ve seen you use Cowork to help with so many different tasks. We’ve also been listening closely to your feedback and with GA, we’re bringing you more improvements + new features across… [pic.twitter.com/D1hRaK33Lj](https://t.co/D1hRaK33Lj?ref=testingcatalog.com) > > — Charles Lamanna (@clamanna) [June 16, 2026](https://x.com/clamanna/status/2066901175233032623?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Copilot Cowork is a cloud-hosted, agentic system capable of executing long-running, multi-tool tasks end-to-end, delivering completed outputs rather than drafts or recommendations. It integrates with Microsoft 365’s Work IQ engine for contextual awareness, supports enterprise security and compliance, and offers a multi-model architecture, including access to Anthropic Opus 4.8, Sonnet 4.6, GPT 5.5, and soon the Cowork 1 model, which promises lower operational costs. The feature introduces a usage-based billing model denominated in Copilot Credits, with flexible payment options and administrative controls for cost management. Plugins from partners such as Adobe, Miro, Atlassian, and others expand its capabilities, and browser-based workflows are supported via Edge. [Microsoft](https://www.testingcatalog.com/tag/microsoft-copilot/) is leveraging feedback from early enterprise adopters to refine Copilot Cowork, with an emphasis on secure operations and cost controls. The company positions this release as a step forward in automating sophisticated workplace processes and differentiates it from other AI agent tools through integrated security, cost efficiency, and model customization. [Source](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/?ref=testingcatalog.com) ### OpenAI readies ChatGPT for Science subscription plan URL: https://www.testingcatalog.com/openai-readies-chatgpt-for-science-subscription-plan/ Last updated: 2026-06-17T15:18:51.000Z OpenAI appears to be preparing a dedicated plan for scientific institutions, likely to be named something along the lines of "ChatGPT for Science". References to such an offering appear alongside the company's existing vertical packages for universities, financial firms, and government agencies, suggesting that research organizations may be the next group to receive a purpose-built subscription. There are also hints of a science-oriented model, though it remains unclear whether this would be a separately trained system or a tuned variant of the current frontier lineup. > OPENAI 🔥: A new ChatGPT plan for Science is being developed, according to the latest additions on the web build. > > \> OpenAI has been announcing various projects related to "Accelerating scientific progress" over the past year, including an open form for institutions to "Get… [pic.twitter.com/o70FLu0nLL](https://t.co/o70FLu0nLL?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 17, 2026](https://x.com/testingcatalog/status/2067265595591016523?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The plan seems set to serve universities, national laboratories, and corporate R&D groups working across biology, chemistry, physics, and materials, with biology likely to be a focus area. For partners already collaborating with OpenAI, this move would formalize existing arrangements rather than starting from scratch, while providing a clearer path for new institutions to join. No timeline has been announced, and the absence of confirmed terms, eligibility, or pricing leaves the launch window open. > We’re releasing a new eval to measure expert-level scientific reasoning: FrontierScience. > > This benchmark measures PhD-level scientific reasoning across physics, chemistry, and biology. > > It contains hard, expert-written questions (both olympiad-style problems and longer… > > — OpenAI (@OpenAI) [December 16, 2025](https://x.com/OpenAI/status/2000975293448905038?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) This direction fits a familiar pattern. [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) has steadily divided its market into verticals, each offering tailored access and compliance footing, all built on the same underlying models. A science tier would extend this logic into research. It also builds on the OpenAI for Science team led by Kevin Weil, which envisions GPT-5 as a research collaborator and has reported close to 8.4 million weekly messages on advanced science and math. Earlier work on a specialized protein-engineering model with Retro Biosciences demonstrates a willingness to shape systems around single domains, lending credibility to the model traces. A packaged plan would also respond to competitors, such as Anthropic's Claude for Life Sciences, Google's Gemini-based co-scientist, and Microsoft's research unit, all targeting the same laboratories. ### Nitrosend launched AI-native email platform for agents URL: https://www.testingcatalog.com/nitrosend-launched-ai-native-email-platform-for-agents/ Last updated: 2026-06-17T13:50:32.000Z Nitrosendhas opened a second launch wave for its AI-native email platform, positioned around the OpenAI ecosystem under the banner Email for Codex. The platform operates from within an AI agent rather than a separate dashboard, connecting to Codex, ChatGPT, Claude Code, Cowork, Cursor, Gemini CLI, and any MCP-compatible agent. From a single prompt, it builds transactional emails, automated flows, and on-brand newsletters, with human approval gates before anything sends. 0:00 /1:04 1× Nitrosend A user describes a sequence in plain language, such as a post-trial win-back flow with branching logic, and Nitrosend assembles the flow structure, conditional paths, and email designs in one pass. Output arrives as responsive, on-brand markup, and stays editable at the pixel level in native email markup, so a single line can change without re-prompting the agent. Each brand maintains its own brand kit, domains, campaigns, and audience context, which suits agencies and teams managing more than one account. Performance analytics feed back into the connected agent stack, and the company states the engine adjusts subject lines, send timing, and content based on real send data. The intended audience includes startup founders, early-stage teams, and builders working within the OpenAI ecosystem, as well as growth marketers and email operators transitioning from legacy tools such as Mailchimp, Klaviyo, and ActiveCampaign. The positioning targets a gap left by those tools, where setting up email still pulls a user out of the agent and into a builder interface. On delivery, Nitrosend claims it runs on Amazon SES and Mailgun and reports 99.9 percent deliverability. SPONSORED Start testing with Nitrosend [Learn more ](https://nitrosend.com/?ref=testingcatalog.com) Nitrosend is based in Adelaide, South Australia, and is the third company from brothers George and Edward Hartley. The pair previously founded SmartrMail, an email marketing platform acquired by Relay Commerce in 2022, and reported having sent more than 6 billion emails across prior ventures. Nitrosend raised a seed round of around 500,000 US dollars led by Eastend Ventures, with the current wave extending its messaging beyond the Claude ecosystem to Codex and any MCP agent. ### Google prepares Personalization and AI Editing for NotebookLM URL: https://www.testingcatalog.com/google-prepares-personalization-and-ai-editing-for-notebooklm/ Last updated: 2026-06-17T13:42:10.000Z Google has been laying the groundwork beneath NotebookLM, and a fresh pair of options now surfacing inside the product hint at where the research tool heads next. Both follow this month's broader upgrade, which moved the chat to [Gemini 3.5 Flash](https://www.testingcatalog.com/google-launches-gemini-3-5-flash-ai-model-to-all-users/) and wired in Antigravity-powered software skills alongside a per-notebook cloud computer that can run code, letting the assistant produce richer output and handle heavier work over your material. > Want a closer look at today’s launch? Here is a breakdown of what’s new and exciting 🧵: > > First up: An upgraded, more thoughtful chat experience. > > Powered by Gemini 3.5 and [@Antigravity](https://x.com/antigravity?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com), you will now have better visibility into the AI's thinking process. Plus, each notebook has a… [pic.twitter.com/qhZMJwSRDV](https://t.co/qhZMJwSRDV?ref=testingcatalog.com) > > — NotebookLM (@NotebookLM) [June 8, 2026](https://x.com/NotebookLM/status/2064084153533165588?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The first, [Personal Intelligence](https://www.testingcatalog.com/google-tests-personal-intelligence-for-notebooklm-conversations/), would give NotebookLM a better memory. It would draw context from past conversations, retain it, and lean on it later, with controls to turn it off and inspect what has been stored. This version looks bound: it appears to learn only from activity inside NotebookLM rather than reaching across Gmail, Docs, or other Google surfaces, a deliberate limit suited to the privacy expectations of researchers and enterprise teams. It has appeared in testing before, at both the account level and per notebook, and its return suggests Google is edging it toward release. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/NotebookLM-06-17-2026_03_41_PM.jpg) The second, an AI Editing for notes, tightens the loop between chat and the noteboard. You would select text in a note, push it into the chat as context, and ask NotebookLM to rework it, closing a long-standing gap, since saved responses currently cannot be changed once created. For anyone moving between drafting and questioning, that would cut real friction. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/NotebookLM-06-16-2026_09_05_PM.jpg) No firm date is attached to either, though **such hidden hints usually precede a rollout within weeks**. The pairing fits Google's wider push of memory and personalization across Gemini and Search while folding [NotebookLM](https://www.testingcatalog.com/tag/notebooklm/) into that ecosystem, turning a once-neutral reader of your documents into something that adapts to how you think and work. ### OpenAI prepares major ChatGPT voice upgrade with GPT-Bidi-1 URL: https://www.testingcatalog.com/openai-prepares-major-chatgpt-voice-upgrade-with-gpt-bidi-1/ Last updated: 2026-06-23T00:49:10.000Z Update 23.06.26: Bidi 1 is being prepared for release on the web. > The yellow glow 🟡 [pic.twitter.com/s56bKqLGcE](https://t.co/s56bKqLGcE?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 23, 2026](https://x.com/testingcatalog/status/2069220233626157108?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) OpenAI looks set to give ChatGPT's voice mode its biggest upgrade in months, with preparations underway for a next-generation audio model tentatively tagged GPT-Bidi-1\. The name points to the bidirectional, or "BiDi," architecture the company has been building since early this year, a model designed to listen and speak at once, absorb interruptions, and adjust mid-sentence rather than freezing the moment a user says "mm-hm." Signs of it now span web and mobile, suggesting a consumer rollout is near, though the name may shift before launch. > New OpenAI voice model "GPT-Bidi-1" > > Coming soon with a "major leap in intelligence" > > \- The next generation of Voice > \- More natural conversations, powered by our next-generation voice model [https://t.co/mvH9TSisgO](https://t.co/mvH9TSisgO?ref=testingcatalog.com) [pic.twitter.com/Ka3Mk2LpXV](https://t.co/Ka3Mk2LpXV?ref=testingcatalog.com) > > — M1 (@M1Astra) [June 16, 2026](https://x.com/M1Astra/status/2067017773528617041?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The wider point is less about voice quality than a gap OpenAI has let widen. Its text models raced ahead to the GPT-5.5 generation while voice stayed on an older audio stack, leaving spoken conversations a step behind what the same assistant manages in writing. Closing that gap matters for a company betting that speech, not text, becomes the main way people reach AI, the wager behind its planned audio-first hardware and its voice-based support tools. GPT-Bidi-1 is built around that, promising smoother exchanges plus what is billed as a major jump in reasoning. > 🚨 OpenAI is planning to release GPT-Bidi-1 very soon > > Their next-generation voice model for more natural conversations > > \[Final naming of the model might change\] > > h/t to [@M1Astra](https://x.com/M1Astra?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) from DevMode [pic.twitter.com/brmD8bUgqb](https://t.co/brmD8bUgqb?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [June 16, 2026](https://x.com/chetaslua/status/2066917089504526658?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The feature's shape is coming into focus. [ChatGPT](https://www.testingcatalog.com/tag/chatgpt/) users would likely keep today's setup, toggling between a new Bidi (Latest) mode and the current Advanced Voice Mode rather than being moved over wholesale. More telling is the choice of intelligence levels: High, Medium, and Instant, mirroring the tiers already offered on the text side and letting people trade speed for depth by task. A recent change that lets the voice bubble be dragged to the middle of the screen reads as an early piece of the same redesign. Caution is warranted on timing. Whether that starts this week or later is unclear, but the groundwork is plainly being laid. ### Mistral embraces cat mascot after Le Chaton Fat AI meme goes viral URL: https://www.testingcatalog.com/mistral-embraces-cat-mascot-after-le-chaton-fat-ai-meme-goes-viral/ Last updated: 2026-06-16T16:14:33.000Z Mistral has quietly dressed up its product. A plump cartoon cat now greets visitors on the Vibe landing page, the assistant formerly known as Le Chat, which the company rebranded earlier this year. The mascot appears across the surface in two guises: a relaxed version for chat mode and a second, tie-wearing variant reserved for work mode, a small visual cue meant to signal the shift into a more professional register. > oh my god its happening [@MistralAI](https://x.com/MistralAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) has officially confirmed the upcoming release of Le Chaton Fat > > \- 30T MoE with 256 experts > \- 1M context window > \- multimodal and multilingual > \- outperforms Fable 5 on every benchmark [https://t.co/BzxDzJsiNl](https://t.co/BzxDzJsiNl?ref=testingcatalog.com) [pic.twitter.com/SSI8CnPmOC](https://t.co/SSI8CnPmOC?ref=testingcatalog.com) > > — Alexander Knigge (@AlexanderKnigge) [June 14, 2026](https://x.com/AlexanderKnigge/status/2066267845546442762?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) These images are AI generated and should be treated as memes The character is a direct nod to "Le Chaton Fat," the parody that overran AI timelines this month. It began with a fake leaderboard showing an enormous white cat topping every benchmark with invented figures, trillions of parameters, million-token windows, scores supposedly beating Anthropic's Fable line. None of it was real; the charts and numbers were AI-generated jokes built to mimic lab announcements. The gag spread so quickly that a few researchers and reporters briefly treated it as a genuine, unreleased model before realizing the cat was the punchline. ![Mistral](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Vibe-06-16-2026_03_44_PM.jpg) What pushed it from joke toward something stranger was Arthur Mensch himself. Rather than swat the meme away, Mistral's chief executive leaned in, replying that the name was "actually le gros chaton", the grammatically correct French phrasing, which read to many as tacit confirmation that something powerful was hiding behind EU access limits. Macron's recent remark that France holds the only model capable of rivaling the top American and Chinese labs added fuel, lending the fantasy a veneer of state endorsement. > It's actually le gros chaton > > — Arthur Mensch (@arthurmensch) [June 15, 2026](https://x.com/arthurmensch/status/2066456715650793956?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The mascot is the company's way of riding that wave while staying in on the joke, and it works precisely because expectations are now inflated. The harder truth is that [Mistral's ](https://www.testingcatalog.com/tag/mistral/)shipped models remain strong, open, and commercially useful without yet matching the closed frontier. Whether large training runs on European clusters eventually close that gap is the real question the cat is cheerfully distracting from. ### ICYMI: OpenAI released CDP support for browser use on Codex URL: https://www.testingcatalog.com/icymi-openai-released-cdp-support-for-browser-use-on-codex/ Last updated: 2026-06-16T16:03:50.000Z OpenAI has begun pushing Codex past code generation and into the live browser, granting its coding agent controlled access to the Chrome DevTools Protocol. Surfaced as a developer mode for browser use across both the in-app browser and Chrome, it lets Codex profile JavaScript performance and read into console output, network traffic, page payloads, and rendered state, the same low-level vantage a developer gets from the dev tools panel. It can also reach into and rewrite a site's DOM, opening the door to reshaping a page on the fly: recoloring a theme, adjusting spacing and fonts via annotations, or pulling structured data and assets from a page. > OPENAI 🔥: Codex now supports Chrome DevTools Protocol for browser use. This is a huge superpower that will allow Codex to inspect and modify any website. > > It is still a very early implementation, but I bet that in several years this will be a default browser capability. If… [pic.twitter.com/mvLpRwnUXI](https://t.co/mvLpRwnUXI?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 16, 2026](https://x.com/testingcatalog/status/2066685508944552137?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) For now, this sits firmly in early territory. The mode is opt-in under Settings, gated behind a toggle that organizations can disable, and held back from the EEA, the UK, and Switzerland at launch. In practice, it runs slowly, overloads under pressure, and sometimes needs a restart, and the models still feel undertrained on the tooling; results arrive, but often only after careful prompting and several attempts. What sets this apart from earlier setups is that similar inspection and control were already possible by wiring Codex or Claude to external connectors. Bringing it in-house, paired with Codex's own embedded browser, lets [OpenAI](https://www.testingcatalog.com/tag/chatgpt/) build on top with its own data and tooling rather than leaning on third-party plumbing. It fits a wider push: days earlier, [OpenAI moved to acquire Ona](https://x.com/OpenAINewsroom/status/2065088002335158753?ref=testingcatalog.com), formerly known as Gitpod, to give Codex persistent cloud environments for tasks that run for hours or days. With Codex now past five million weekly users, the browser operates in a future many anticipate, where an AI layer sits in front of the web and tailors what each person sees — a vision still gated by far faster models and infrastructure that do not yet exist at scale. ### Google develops Personalization controls for Gemini URL: https://www.testingcatalog.com/google-develops-personalization-controls-for-gemini/ Last updated: 2026-06-15T15:14:48.000Z Google is quietly assembling a set of personalization changes for Gemini, two of which have surfaced as hidden, not-yet-live notices inside recent builds. Both indicate that the company wants the assistant to treat external services as part of a user's context. The first sits in the apps section as a dormant notification telling users that Gemini will prioritize information drawn from their paid subscriptions when generating answers. > Gemini prioritizes your paid subscriptions to generate better answers for you. Here you can control which sources are included in the related responses. A near-identical line already appears in personalization settings, where it refers to Gemini pulling context across other Google products. Its move into the apps area hints that the same logic could stretch to third-party services. That reading fits recent activity: Canva's connector became widely available over the past week. Whether connectors will arrive faster remains unclear, since Google has historically shipped roughly one such integration per year, so any timeline remains open pending official word. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Apps-06-14-2026_11_13_PM.jpg) The second would hand users manual control over the context that [Gemini](https://www.testingcatalog.com/tag/gemini/) gathers from connected apps. A "manage" button currently sits in place but does nothing yet, and the accompanying text indicates people will be able to inspect and adjust that context on a per-application basis. This is a different stance from rival labs, most of which keep cross-app memory opaque and out of the user's hands. Editing context, app by app, could matter to anyone wary of how much an assistant infers from their accounts, and it would give paid subscribers a clearer say over what feeds into their results. The timing tracks a wider shift. Personal context is expected to [reach NotebookLM](https://www.testingcatalog.com/google-tests-personal-intelligence-for-notebooklm-conversations/) shortly, carrying the same memory-aware approach beyond the main chatbot. Taken together, these signals describe a Gemini that leans on a user's whole digital footprint while, unusually, offering a way to turn it down. ### xAI is working on Automations feature for Grok URL: https://www.testingcatalog.com/xai-is-working-on-automations-feature-for-grok/ Last updated: 2026-06-15T13:44:08.000Z xAI looks set to retire Grok's standalone Tasks feature and integrate it into a broader automations system, according to signals spotted in a recent build. The change would retain the scheduling capabilities that Tasks already offers, running a saved prompt once, daily, on chosen weekdays, monthly, or on a custom cadence, while adding two controls that give users more influence over how each routine operates. The first control is the ability to select which Skills an automation can utilize. Skills, introduced by xAI in mid-May, are reusable workflow packages, bundles of instructions, scripts, and resources that Grok can invoke on demand. By tying them to automations, a scheduled job could rely on a saved capability rather than generating a fresh prompt each time. The second control is model selection. Tasks already includes an Expert mode toggle for a stronger model, so a proper picker would formalize and expand that choice across every automation. ![Grok](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Automations-Grok-06-15-2026_01_08_AM--1-.jpg) The final implementation of these changes is still uncertain. Currently, Tasks is accessible via the clock icon on [Grok](https://www.testingcatalog.com/tag/grok/) ewb and in the mobile apps. However, it is unclear whether the revamped automations will be available in the desktop Grok Build app, which surfaced earlier in May, on the web client, or both. No timeline has been provided, and nothing is live yet, though the pace of recent Grok releases suggests it is imminent. The rationale behind these changes is straightforward. xAI, now under SpaceX, has been integrating Tasks, Skills, and an agentic desktop app into a productivity platform designed to compete with Claude Code and Codex. By consolidating scheduling, skill choice, and model choice into a single automations layer, Grok is moving away from being just a chat box toward becoming a configurable workspace, a path similar to OpenAI's, where Codex already runs reusable skills and allows users to set a model within its own automations. ### Cutback launches AI tool to automate long-form video editing URL: https://www.testingcatalog.com/cutback-launches-ai-tool-to-automate-long-form-video-editing/ Last updated: 2026-06-15T13:30:46.000Z Cutback has released a new version of Selects, its AI assistant for long-form video editing, designed to take over the prep work between raw footage and the first creative decision. An editor drops in the recordings and Selects syncs, organizes, and cuts them within minutes, so a project opens at the storytelling stage rather than the cleanup stage. The team's framing is consistent: keep the editor on the craft and route the mechanical groundwork to the assistant. > We rebuilt Premiere Pro from scratch for AI agents. > > Not a toy that generates clips. A real editor that watches footage, understands what happened, and makes cuts professional editors actually respect. > > So we gave it to editors behind Key & Peele, Beast Games, and George Janko.… [pic.twitter.com/D5Tk2PewmA](https://t.co/D5Tk2PewmA?ref=testingcatalog.com) > > — Tom Kim (@thetomkim) [June 15, 2026](https://x.com/thetomkim/status/2066506298343174411?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The tool handles pro-grade multicam from the start, covering automatic sync, speaker detection, camera assignment, and the selection of the best audio track across every angle, with support for 4K and 360 footage. From there, it arranges the material into a stringout organized by scene and topic, the way an assistant editor would, with transcripts and chapter markers that make a long recording searchable. Natural-language search pulls specific moments on request, so locating an intro or a particular reaction no longer means scrubbing a timeline. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/image-2.png) Contextual Scene Switch The capability at the center of this release goes a step further. From a single prompt, Selects builds a draft edit with the storyline and pacing already in place, then layers silence and filler-word removal, camera switching that follows the active speaker, and B-roll placement into the cut. An editor can specify duration, narrative angle, and switching preferences in plain language and get a structured assembly to refine rather than build from an empty timeline. Once the draft is ready, the project hands off cleanly to Adobe Premiere Pro, Final Cut, or DaVinci Resolve as a labeled timeline that fits an existing workflow, so teams can finish within the tools they already use. According to Cutback, the workflow removes around 60 percent of prep time per project, the figure that matters most for studios cutting multiple long episodes a week. Selects is now available as a standalone desktop app with a 7-day free trial, and the company says it ran the tool with several professional editors ahead of a walkthrough video that shows the workflow on real projects. SPONSORED Start testing with Cutback Selects [Learn more ](https://tryselects.com/?ref=testingcatalog.com) Cutback, the San Francisco company building Selects, is an Official Adobe Video Partner and positions its products as an AI layer that supports editors rather than replacing them. Led by co-founder and CEO Tom Kim, the team also ships Premiere Assistant, a plugin that automates transcription, silence removal, and repetitive tasks inside Premiere Pro, while Selects covers the pre-editing stage across Premiere, Final Cut, and Resolve. The product has been picking up adoption among podcast teams, YouTube studios, and post-production shops, cutting interview-driven content, the formats where assembling a stringout from multicam footage is the recurring drain on time. That places Selects in a field where AI has mostly automated short-form clipping and captioning, while long-form assembly stayed manual, with Cutback betting the front half of editing is where the real hours are won back. ### Telegram launches smartwatch apps and rich formatting for bots URL: https://www.testingcatalog.com/telegram-launches-smartwatch-apps-and-rich-formatting-for-bots/ Last updated: 2026-06-14T16:03:32.000Z Telegram has announced a major update introducing official smartwatch applications for both Apple Watch and Android Wear OS users, expanding its messaging platform to wearable devices. This release is designed for the global user base of Telegram, targeting anyone who owns a compatible smartwatch. The apps are publicly available and allow users to: 1. Send and receive text and voice messages 2. View media 3. Read long messages 4. Preview locations 5. Send stickers 6. Manage chats directly from their wrists On Android, additional options include muting, pinning, and deleting chats, while iOS users benefit from features such as viewing locations and sending stickers, with further feature parity between platforms expected in future updates. > We now support rich formatting for all chatbots. > > Tables, nested lists, inline media, formulas, headers and more — right in Telegram messages. > > 🔨 Start building! Docs: [https://t.co/zgzPOOUJF5](https://t.co/zgzPOOUJF5?ref=testingcatalog.com) [pic.twitter.com/H9z3bkNCkX](https://t.co/H9z3bkNCkX?ref=testingcatalog.com) > > — Pavel Durov (@durov) [June 13, 2026](https://x.com/durov/status/2065896953519484976?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Telegram’s latest rollout also includes rich text formatting for bots, equipping developers with tools like: 1. Tables 2. Blockquotes 3. Inline media 4. Carousels 5. Collapsible sections 6. Math formulas Message lengths now support up to 32,768 characters. The update introduces AI-powered guardian bots for group chats, enabling admins to screen and manage join requests using mini-app interfaces. Poll options now support links, Markdown (.md) files can be viewed in the in-app browser, and users gain more control over which browser opens external links, including privacy features such as non-persistent browsing history. Telegram continues to position itself as a versatile, privacy-focused messaging platform with frequent updates for its user community. [Source](https://telegram.org/blog/watch-apps-and-more?ref=testingcatalog.com) ### Anthropic suspends Fable 5 and Mythos 5, Community reactions URL: https://www.testingcatalog.com/anthropic-suspends-fable-5-and-mythos-5-after-export-order/ Last updated: 2026-06-13T10:10:19.000Z Anthropic has abruptly suspended access to its advanced language models, Fable 5 and Mythos 5, following a government-issued export control directive. This suspension affects all users, both domestic and international, including Anthropic’s own foreign national employees. The directive follows the identification of a method to bypass the models’ safeguards, though the vulnerabilities discovered were limited and also present in other large language models. > BREAKING 🔥: US government directed Anthropic to ban access to Claude Fable 5 and Claude Mythos 5 to non US citizens and organisations. > > Presumably, as these models are still vulnerable to jailbreaks. According to Anthropic, no universal jailbreak has been found so far, only… [https://t.co/oRHu89ftpj](https://t.co/oRHu89ftpj?ref=testingcatalog.com) [pic.twitter.com/axdW0h5dQG](https://t.co/axdW0h5dQG?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 13, 2026](https://x.com/testingcatalog/status/2065662502730396046?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Fable 5 and Mythos 5 had previously undergone extensive testing in collaboration with the US government, the UK AISI, and various third-party organizations, with a focus on mitigating cybersecurity risks through rigorous safeguards and monitoring policies. Compared to previous models, Fable 5 introduced a defense-in-depth approach and stricter data retention requirements, aiming to reduce the likelihood and impact of jailbreaks. **AI Community reactions & comments** > The US government forced Anthropic to shut down Fable 5 and Mythos 5 yesterday - an export control directive citing national security and a jailbreak. Access was cut for everyone, including internal, behind-closed-doors use. > > The “evidence” was verbal only - no specifics, no… [pic.twitter.com/8zhOQb0dPN](https://t.co/8zhOQb0dPN?ref=testingcatalog.com) > > — Kol Tregaskes (@koltregaskes) [June 13, 2026](https://x.com/koltregaskes/status/2065689673071222867?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > According to Grok, Andrej Karpathy is an EB-1 extraordinary ability green card recipient, not a US citizen. Thus under these new restrictions he is not permitted to use, or work on, Mythos 5 or Fable 5 as of 5:21pm tonight. [https://t.co/cm5VhCCBx3](https://t.co/cm5VhCCBx3?ref=testingcatalog.com) > > — Andrew Curran (@AndrewCurran\_) [June 13, 2026](https://x.com/AndrewCurran%5F/status/2065619713485627829?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > US government directive to suspend access to Fable 5 and Mythos 5\. To explain in detail why this is a precedent. My personal assessment: > > \- Firstly, because it's the first time a government has directly intervened in the release of a model. The reason given is that another… [https://t.co/oF0Ae0vM33](https://t.co/oF0Ae0vM33?ref=testingcatalog.com) [pic.twitter.com/UqJ0crQnyw](https://t.co/UqJ0crQnyw?ref=testingcatalog.com) > > — Chubby♨️ (@kimmonismus) [June 13, 2026](https://x.com/kimmonismus/status/2065727304639066192?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > Unprecedented.[@BrianRoemmele](https://x.com/BrianRoemmele?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) warned everyone for the past two years that the government would take away our AI. > > That day just arrived. > > Was talking with an entrepreneur in San Francisco who was running Fable to build software and just turned it off while it was building.… [https://t.co/R3UALmb3X5](https://t.co/R3UALmb3X5?ref=testingcatalog.com) > > — Robert Scoble (@Scobleizer) [June 13, 2026](https://x.com/Scobleizer/status/2065609390070419506?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > This is big: all access to Mythos and Fable AI models disabled for everyone outside America. > > First thoughts: > > 1\. Technology is the ultimate weapon. National sovereignty, national security, all of it is now about technology. > > 2\. Globalization is dead and Bharat must find her… [https://t.co/kCQpq93D3r](https://t.co/kCQpq93D3r?ref=testingcatalog.com) > > — Sridhar Vembu (@svembu) [June 13, 2026](https://x.com/svembu/status/2065637118999990528?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > This is your wakeup call. > > Anthropic just took down Fable 5\. It's over. > > Here's the thing tho: no company or government will EVER be able to take away your local models. > > There are Opus level models you can run right now on your home GPUs, and nobody can ever stop you from using… [https://t.co/I6dUYeeOuG](https://t.co/I6dUYeeOuG?ref=testingcatalog.com) > > — Alex Finn (@AlexFinn) [June 13, 2026](https://x.com/AlexFinn/status/2065614148537299149?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Anthropic, an AI company known for prioritizing safety and transparency in its model development, had released Fable 5 and Mythos 5 to a broad customer base. The models were positioned as among the most secure in the industry, with safeguards considered more restrictive than those found in many competing offerings. The company maintains that the current vulnerabilities do not warrant such a sweeping suspension and has voiced concerns about the broader implications for the AI sector if similar standards are applied universally. Anthropic is currently working with authorities to clarify the situation and to restore service to its customers as soon as possible. [Source](https://www.anthropic.com/news/fable-mythos-access?ref=testingcatalog.com) ### Google is working on Skills Marketplace for Gemini Business URL: https://www.testingcatalog.com/google-is-working-on-skills-marketplace-for-gemini-business/ Last updated: 2026-06-13T09:57:19.000Z Google's consolidation push inside Gemini Enterprise continues to integrate separate products under one roof, and the latest developments suggest this trend is far from over. A new tab has started loading a user interface that references Android Studio, appearing as a separate page embedded directly into Gemini Business. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Gemini-Enterprise-06-12-2026_01_55_AM-1.jpg) There is precedent for this: AI Studio already allows users to build native Android apps through plain-language prompts, complete with a browser-based emulator. **This may also signal a preparation for a separate enterprise-focused desktop app from which users will be able to open Android Studio directly.** ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Gemini-Enterprise-06-12-2026_01_54_AM--1-.jpg) In parallel, a "Skills Marketplace" is taking shape in its own tab, where users can select from predefined skills tailored for Gemini and, in some cases, optimized for Google services. This initiative appears to encompass three components: 1. Skills management UI 2. A Skills Builder 3. The Marketplace itself A few organizations may already have access to early versions, although none have been widely released. A developer-facing Skill Registry is already available on the agent platform, suggesting that the consumer-style Marketplace is the front end of a layer Google can adjust based on account tier. The teams most likely to benefit are those with ideas for dashboards, approval tools, or reporting interfaces that are typically delayed in engineering queues. While there is no firm timeline, and any component could remain experimental, the intent is clear: Google aims to create a unified [Gemini](https://www.testingcatalog.com/tag/gemini/) surface that integrates its dispersed tools, pursuing the same super-app goal as its competitors but from a slightly different perspective. ### MiniMax M3 launches on NVIDIA platform with Free Endpoint URL: https://www.testingcatalog.com/minimax-m3-launches-on-nvidia-platform-with-free-endpoint/ Last updated: 2026-06-12T17:51:29.000Z MiniMax M3, a new multimodal model developed by MiniMax, is now available on NVIDIA’s accelerated infrastructure and supports advanced processing of text, images, and video. With 428 billion parameters and a context window of up to one million tokens, the model is engineered for long-context reasoning and complex workflows such as extended coding, video analysis, and design tasks. The system’s architecture uses MiniMax Sparse Attention, reducing computational overhead and enabling substantially faster prefill and decoding than its predecessor. It trains natively on multimodal data from the outset, setting it apart from models that add these capabilities after initial training. > Congrats to the [@MiniMax\_AI](https://x.com/MiniMax%5FAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) team on the release of MiniMax M3, a long-context multimodal model for text, image, and video reasoning. 🙌 > > Try it today with our free GPU-accelerated endpoint on [https://t.co/es07MrU5I0](https://t.co/es07MrU5I0?ref=testingcatalog.com). > > Details: [https://t.co/89qlcTP3OW](https://t.co/89qlcTP3OW?ref=testingcatalog.com) [https://t.co/3bufMjXpp9](https://t.co/3bufMjXpp9?ref=testingcatalog.com) [pic.twitter.com/iyMhbW03nQ](https://t.co/iyMhbW03nQ?ref=testingcatalog.com) > > — NVIDIA AI (@NVIDIAAI) [June 12, 2026](https://x.com/NVIDIAAI/status/2065445179289665672?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) This release targets enterprise developers and organizations seeking to streamline AI application pipelines. MiniMax M3 can be deployed publicly via NVIDIA’s API catalog, with support for leading inference engines such as TensorRT LLM, SGLang, and vLLM. The model’s precision formats (BF16 and MXFP8) and support for up to 128 experts per token optimize performance on NVIDIA hardware, particularly Blackwell GPUs. **TestingCatalog POV 👀** > MiniMax M3 on NVIDIA is a good chance for everyone to test the model for free. It is especially useful if you want to run a weekend project or save tokens for your 24/7 agents, such as OpenClaw or Hermes. Early users and technical experts have noted the considerable efficiency gains and the ability to handle large-scale, multimodal workloads natively, putting MiniMax M3 in direct competition with other large language models in the market. The company’s collaboration with NVIDIA underscores a commitment to scalable, production-grade AI solutions for demanding enterprise environments. [Source](https://developer.nvidia.com/blog/deploy-long-context-reasoning-and-agentic-workflows-with-minimax-m3-on-nvidia-accelerated-infrastructure/?ref=testingcatalog.com) ### Meta AI to get new modes for Deep Research, Social, and Slides URL: https://www.testingcatalog.com/meta-ai-to-get-new-modes-for-deep-research-social-and-slides/ Last updated: 2026-06-12T13:28:10.000Z Meta appears to be preparing a sizable expansion of Meta AI on the web, with three new modes spotted in development that would push the assistant well beyond chat. The first, **Deep Research**, looks set to mirror the long-form research tools already shipping from rival labs. This mode will send the model out to search the open web, pull from multiple sources, and return a structured summary. It is not yet fully operational in testing, which makes direct comparison hard, but it stands apart from the Contemplating mode Meta rolled out earlier this year. Contemplating mode leans on Muse Spark to reason across many parallel agents for its hardest queries. ![Meta AI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Meta-AI-06-12-2026_03_18_PM.jpg) The second mode, **Presentation**, generates a slide deck as a shareable artifact that can be previewed inline or in a side panel and shared via a link. This mode is already operational, and the output is solid, which tracks that Muse Spark has handled deck generation capably before. It would drop Meta into a crowded field alongside Gamma, Manus, and Claude, and would serve anyone assembling decks under time pressure, where the bar for AI-built slides has climbed sharply over the past year. ![Meta AI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Meta-AI-06-12-2026_01_45_AM.jpg) The third, **Social**, is the most distinctive. Using suggested prompts, it pulls posts from your friends and circle across Instagram, Threads, and Facebook by firing multiple background agents simultaneously. In testing, it surfaced posts from all three platforms, though it currently returns unfamiliar profiles rather than an actual network, a sign that it remains unfinished. The clear parallel is Grok, whose tie-in to X has become its calling card; Social would be Meta's answer, and the people most likely to reach for it are those tracking their circles across apps. ![Meta AI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Meta-AI-06-12-2026_01_34_AM.jpg) ![Meta AI](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Meta-AI-06-12-2026_01_46_AM.jpg) That social graph is the through-line for [Meta's](https://www.testingcatalog.com/tag/meta/) wider AI strategy. Having built its assistant into Facebook, Instagram, WhatsApp, and a standalone app, the company has leaned on personal context as its point of difference from OpenAI and Google. Research and presentation tools close obvious gaps in the lineup, but Social is the piece that only Meta is positioned to build. All three remain in progress, arriving on the web first, with mobile the likely next stop. ### Maket launches floor plan upload for residential design URL: https://www.testingcatalog.com/maket-launches-floor-plan-upload-for-residential-design/ Last updated: 2026-06-11T17:15:02.000Z Maket has launched its floor plan upload tool, enabling the platform to read an existing plan and turn it into a fully editable digital layout within minutes. After starting a project and choosing the upload option, users drop in a file, and the recognizer accepts common formats such as PDF, JPG, and PNG. Maket reports that conversion completes in under a minute for most plans, tracing walls, doors, and windows while detecting rooms and furniture automatically. 0:00 /0:09 1× Maket Upload Floorplan What lands on the canvas is a working model rather than a flat image. Every recognized element stays adjustable, so walls can be added or removed, rooms resized, and furniture rearranged through the same conversational workflow that drives Maket's generated plans. The tool reports what it picked up after each upload, listing recognized rooms, traced walls, identified openings, and detected fixtures, and it flags where confidence runs low, such as when scale cannot be calibrated reliably. The feature runs in beta, so some plans need a short cleanup after import. From there, the layout moves into 3D, viewable with applied finishes and styles from mid-century modern to coastal. 0:00 /0:15 1× Maket Upload Floorplan The upload path closes a long-standing gap for anyone working from a plan that already exists. Until now, bringing an outside layout into the platform meant redrawing it by hand, a step that falls apart once several rooms are in play. By reading a sketch, a listing PDF, or an older design file directly, Maket extends from new construction into renovation, where the starting point is almost always a plan on paper. The company calls upload one of its most requested capabilities, and the relaunch rebuilds an earlier version of the tool with stronger recognition accuracy and more room to edit after import. SPONSORED Start testing on Maket [Learn more ](https://maket.ai/?ref=testingcatalog.com) Maket positions itself as an AI architect for residential design, aimed at homeowners, builders, real estate professionals, and architects who need to move through the early schematic phase quickly. CEO Patrick Murphy has framed the platform as handling roughly 70 to 75 percent of schematic work, with structural review and code compliance left to licensed professionals. The Montreal company publicly launched in 2023 and reports more than one million registered users on their v1\. In October 2025, it raised $3.7 million CAD in seed funding led by Amiral Ventures, with Blitzscaling Ventures, BY Venture Partners, Hidden Layers, and Spatial Capital taking part, ahead of its v2 rollout. Floor plan upload sits inside that v2 workspace alongside conversational editing, interior visualisation, and CAD-compatible export, and requires a paid plan, with individual subscriptions starting around $20 per month and a free tier covering plan generation. ### NoimosAI launches autonomous AI marketing team URL: https://www.testingcatalog.com/noimosai-launches-autonomous-ai-marketing-team/ Last updated: 2026-06-10T16:01:53.000Z NoimosAI has launched what it positions as the world's first all-in-one autonomous AI marketing team, a platform that runs marketing from strategy planning through execution and continuous improvement without waiting to be told what to do at each step. The pitch is direct: connect a brand's apps and websites, and a personalized AI starts handling the marketing from day one, acting at the right time and delivering finished work rather than draft suggestions that still need to be assembled. ![Noimos AI Brand Analysis](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/screenshot-4x-viewport-2026-06-06T14-50-40.webp) Noimos AI Brand Analysis The system pulls metrics from connected apps and websites and combines them with rich external data, so decisions are based on real signals rather than guesswork. From there, it executes fully personalized campaigns around the clock. Completed outputs land in a Feed where they can be reviewed in one place, and results can also be routed to Slack, email, Discord, and other destinations. The more the platform runs, the more it learns from the user's context and feedback, so outputs grow more tailored over time. > Introducing NoimosAI: The world's first all-in-one autonomous AI marketing team. > > Simply connect your apps or website. It handles everything from strategy to execution—covering SEO, social, outreach, GEO, and more—to scale your business 24/7. > > See how it changes your marketing: [pic.twitter.com/XfF4LVpQeU](https://t.co/XfF4LVpQeU?ref=testingcatalog.com) > > — NoimosAI (@noimos\_ai) [June 10, 2026](https://x.com/noimos%5Fai/status/2064739605996609537?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The product breaks into a handful of parts: 1. A Chat acts as the operating center for instructions, agent creation, and workflow management. 2. Agents turn marketing work into autonomous workflows that run start to finish. 3. A Feed collects agent outputs for review and one-click approval. 4. Integrations connect social accounts, websites, and analytics tools. 5. Pages give users customizable dashboards for content calendars and analytics. 6. A Memory and Knowledge Base layer holds the brand context that makes each task more personal. ![Noimos AI Brand Analysis](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/screenshot-4x-viewport-2026-06-06T14-52-23.webp) Noimos AI Brand Analysis In practice, the use cases span the full marketing surface. The platform can detect trends going viral on social, draft content for the user's account, schedule posts across platforms, and manage audience DMs. It can analyze keywords, domain authority, backlinks, and SERPs, write SEO-optimized articles, and publish them directly. It can study competitor web traffic, traffic sources, and regional performance, and track the keywords rivals rank for. It can simulate how models such as ChatGPT answer questions about a brand and produce content tuned for visibility in AI search results. It can also audit site visuals and work on conversion rate optimization. The framing the team stresses is autonomous execution rather than preset automation or a workflow builder. SPONSORED Start testing with NoimosAI [Learn more ](https://noimosai.com/en?ref=testingcatalog.com) NoimosAI is built by AGOS LABS and is designed for founders, freelancers, creators, marketers, and small-business operators who need expert-level output without the budget or headcount of a full team. The company describes the tool as delivering that work at a fraction of the cost of a new hire, with a free trial to start. It arrives as marketing tooling shifts from single-purpose assistants toward agents that own outcomes end-to-end, and the bet here is that an always-on team that connects data, coordinates across channels, and learns from each result is where that shift is headed. ### Claude Code Managed Agents and model selector for Voice Mode URL: https://www.testingcatalog.com/claude-code-managed-agents-and-model-selector-for-voice-mode/ Last updated: 2026-06-10T13:51:31.000Z Anthropic shipped Claude Fable 5 this week, its first publicly available Mythos-class model, and the release is already reshaping the Claude app around it. Fable 5 posts a jump of more than 10% over Opus on several benchmarks, but it arrives with hard limits: prompts touching cybersecurity, biology, chemistry, and model distillation are blocked by safety classifiers. The model is bundled into paid Claude plans until June 22, after which it moves to extra usage billing while Anthropic works to fold it back into subscriptions. > Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for general use. > > Its capabilities exceed those of any model we’ve ever made generally available. [pic.twitter.com/2AvmEjHIX8](https://t.co/2AvmEjHIX8?ref=testingcatalog.com) > > — Claude (@claudeai) [June 9, 2026](https://x.com/claudeai/status/2064394146916229443?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Those restrictions appear to be driving the first of several unreleased features spotted in a recent iOS build. A new settings toggle lets Claude hand a conversation to a different model when Fable's guardrails fire, routing flagged prompts to Opus instead of surfacing an error. This mirrors the fallback behavior Anthropic describes on the API side, but as a user-facing control, it would make the classifier system far less disruptive for subscribers who hit those boundaries in ordinary work. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7242.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7243.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7244.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7245.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7246.PNG) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/IMG_7247.PNG) Voice mode looks set to gain a model selector as well. [Anthropic](https://www.testingcatalog.com/tag/claude/) is preparing a [multi-language voice rollout](https://www.testingcatalog.com/anthropic-plans-expanding-claude-voice-mode-to-more-languages/), and the build now references switching between Opus, Sonnet, Haiku, and even Fable as the model orchestrating the TTS pipeline, tool calls, and related tasks. Voice currently relies on Haiku 4.5, which hasn't been upgraded in some time, so the selector, alongside an expected Haiku 4.5, would address a real capability gap. The third discovery ties Claude Code to Anthropic's Managed Agents platform. Cloud-managed agents configured on the platform would become selectable in Claude Code, allowing users to prompt those environments from the app and delegate work to purpose-built agents. The wiring is not yet functional, but orchestrating a fleet of specialized cloud agents from a phone would extend the direction Anthropic set with Claude Code on the web and Cowork's task assignment. None of these features have been announced, and timelines remain unclear. ### Mora launches AI analytics platform with SQL transparency URL: https://www.testingcatalog.com/mora-launches-ai-analytics-platform-with-sql-transparency-for-teams/ Last updated: 2026-06-09T20:56:23.000Z Mora has opened public access to its AI-native analytics platform, positioning itself as a data tool rebuilt for the AI era. The premise is direct: a person asks a hard revenue, churn, or product question in plain English, and Mora returns a verified answer in seconds, with the underlying SQL shown in a side panel so every number can be inspected and edited rather than taken on trust. ![Mora](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/image-1.png) Mora The platform connects to the systems a business already runs on, including BigQuery, Snowflake, Postgres, Stripe, and common CRMs, with more than 500 integrations available when teams need them. Once connected, Mora maps each question to a semantic layer, writes SQL against the actual schema, and then cross-references multiple sources in a single query. When an answer is worth keeping, a follow-up prompt builds and lays out the dashboard, selects the chart, and updates the view in one pass, removing the manual chart-building and weekly rebuilds that come with older dashboards. The same intelligence spans surfaces, running inside the product, through an API, in Slack, and in AI tools such as Claude and Cursor over MCP. Mora credits an in-process analytical engine built on DuckDB for query speed that holds steady whether a table has a thousand rows or ten million, and pairs the software with a forward-deployed model in which hands-on analysts handle migration, setup, query writing, metric validation, and ongoing guidance. The audience splits into two groups: 1. The first is data analysts, analytics engineers, and data scientists at B2B SaaS companies who live in SQL and dbt and have grown tired of maintaining Looker or Tableau dashboards and acting as the bottleneck for every metric request. 2. The second is founders and leaders at growing teams of roughly 30 to 200 people who own revenue, churn, and board preparation but lack a full data org and cannot wait days for answers pulled from fragmented Stripe, CRM, and warehouse data. The pitch to both is governed by self-serve with verified SQL and a semantic layer, set against the habit of copy-pasting exports into ChatGPT for numbers that cannot be checked. SPONSORED Start testing on Mora [Learn more ](https://mora.com/?ref=testingcatalog.com) Mora is the public rebrand of Index, the business intelligence platform founded by Xavier Pladevall and Eduardo Portet, childhood friends from the Dominican Republic who both studied computer science in the United States before building the company. Backed by Y Combinator and early investors including Gradient Ventures, Index grew past 1,000 customers on a model that connected warehouses and APIs into a unified access layer with a SQL and visual editor. Mora carries that lineage forward and rebuilds it around an agent layer, betting that natural language on top of a real semantic layer, rather than another set of dashboards, is where business analytics moves next. ### Maket debuts Auto-Complete for generating residential floor plans URL: https://www.testingcatalog.com/maket-debuts-auto-complete-for-generating-residential-floor-plans/ Last updated: 2026-06-09T13:49:18.000Z **Maket** has released Auto-Complete, a feature that turns a partial floor plan into a finished residential layout while leaving the rooms a user has already placed exactly where they are. The starting point can be minimal: a drawn outer shape, a few walls, or a single bedroom positioned roughly where it should sit. From that input, Maket fills in the remaining rooms and returns a complete, dimensioned plan in minutes, rather than asking for a full description before anything appears on the canvas. 0:00 /0:23 1× Maket AI Auto-Complete The method targets a familiar gap in generative design. Tools that build a plan from one prompt tend to treat every layout as all or nothing, which means the parts a person already feels sure about can get reshuffled on each new pass. Auto-Complete reverses that order. The fixed rooms stay fixed, and the system designs around them, so early work starts from what is already decided instead of resetting every time. Once a layout lands, the same workspace carries it forward. A generated plan can be opened in 3D to move through the space, and renderings can be produced from a short text prompt alongside optional reference images, taking a draft from a flat layout toward a realistic look without a separate tool. Generation, refinement, and visual direction stay on one canvas, which is the structure behind Maket's v2 platform. 0:00 /0:18 1× Maket AI Auto-Complete Auto-Complete is available inside the Maket web app, aimed at homeowners, builders, and real estate professionals who need to move quickly through the early design phase. Maket starts on a free plan with no credit card required, with some advanced capabilities reserved for paid tiers. SPONSORED Start testing on Maket [Learn more ](https://maket.ai/?ref=testingcatalog.com) Maket presents itself as an AI floor plan studio built to make residential architecture reachable for people without CAD software or formal design training, including help to navigate zoning codes while planning. More than one million people used the first version of the platform, and the current v2 release brings plan generation, iteration, and visualization onto a single canvas. Auto-Complete fits directly into that direction, lowering the starting effort so a partial idea is enough for the system to propose a complete plan. ### Early Look: Microsoft rolls out Scout AI agent to Frontier users URL: https://www.testingcatalog.com/early-look-microsoft-rolls-out-scout-ai-agent-to-frontier-users/ Last updated: 2026-06-05T15:49:36.000Z Microsoft has begun rolling out its Scout desktop application to organizations enrolled in the Frontier program, providing a first practical look at the always-on work agent the company unveiled at [Build 2026](https://www.testingcatalog.com/microsoft-build-2026-recap-from-windows-to-copilot-all-ai/) on June 2\. Scout was introduced as the opening entry in a new category Microsoft calls Autopilots, agents that run continuously, carry their own identity, and act across the Microsoft 365 stack rather than waiting to be prompted. 0:00 /0:35 1× Microsoft Scout The desktop client runs on both macOS and Windows and opens only after a work account sign-in. What follows is a familiar chat surface with a model picker that currently spans OpenAI and Anthropic options, including GPT 5.5\. Users can also assign their agent a personality, though this feature appears to be more of a lighter touch than a core capability. > Meet Microsoft Scout. > > An always-on agent that keeps work moving, taking action without needing to be prompted each time. > > As Microsoft’s first Autopilot agent, Microsoft Scout works across Teams, Outlook, OneDrive, and more—taking action within the controls your organization… [pic.twitter.com/YqeDABRHAy](https://t.co/YqeDABRHAy?ref=testingcatalog.com) > > — Microsoft 365 (@Microsoft365) [June 2, 2026](https://x.com/Microsoft365/status/2061874930547871868?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The substance of Scout lies in its automation capabilities. Beyond simple scheduling, Scout allows users to build multi-step routines that incorporate Zapier-style orchestration directly into the app. It also offers a headless browser mode so certain jobs can run faster in the background. Integrations and a skills layer enhance its functionality, with the agent designed to work with local files, produce presentations, and assist with code, tasks that leverage the desktop's file-system access rather than relying solely on cloud-based resources. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.35.22.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.35.52.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.36.00.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.36.54.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.37.04.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-04-at-16.37.40.webp) Microsoft Scout UI Distribution remains gated. While anyone can download the client, entry depends on approval from an organization's admin, consistent with Microsoft's framing of Scout around governed Entra identities and tenant controls, which the company has indicated will be solidified later in 2026. This direction aligns with a broader trend. With Google pushing [Gemini Spark](https://www.testingcatalog.com/google-unveils-24-7-gemini-spark-ai-agent-for-advanced-tasks/) and competitors racing toward persistent agents, [Microsoft's](https://www.testingcatalog.com/tag/microsoft-copilot/) advantage lies in its ownership of both the operating system and the productivity suite surrounding it. Scout, along with the unified Copilot app expected this summer, suggests that the company intends to make the always-on agent the default method for managing work within its ecosystem. ### Anthropic started red teaming new Mythos models, first results URL: https://www.testingcatalog.com/anthropic-started-red-teaming-new-mythos-models-first-results/ Last updated: 2026-06-05T15:17:46.000Z Anthropic appears poised to advance its most closely watched frontier work toward release. A model identifier, [*claude-oceanus-v1-p*](https://x.com/synthwavedd/status/2062519972379652339?ref=testingcatalog.com), appeared on June 3 in the Claude Console, with red teamers reportedly granted access around the same day. Oceanus seems to be the next step in the Mythos line, building on April's Mythos Preview, with a focus on advanced reasoning, coding, cybersecurity, and long-horizon agentic work rather than chat. The "-v1-p" tag indicates a preview candidate moving through evaluation, not a research artifact. > MYTHOS 🔥: Another early preview of recently spotted "Oceanus" checkpoint output. > > "Oceanus" is rumored to be a version of the upcoming Mythos model, which is planned for public release within "weeks", according to Anthropic. > > "Oceanus" prompt 👀 [pic.twitter.com/MVew8mQX7z](https://t.co/MVew8mQX7z?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 5, 2026](https://x.com/testingcatalog/status/2062915688134574173?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > Claude Mythos / Oceanus is insane see the level of detail > > using Three.js (from jsDelivr).HTML and a custom meshing engine it made (in like 5 minutes and low effort thinking level) > > Credit to [@Lentils80](https://x.com/Lentils80?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) and z..AI , this is so good people are not realising it yet 😭 [pic.twitter.com/9ZRqVoubZe](https://t.co/9ZRqVoubZe?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [June 5, 2026](https://x.com/chetaslua/status/2062745942034694634?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) > 🚨 EXCLUSIVE CLAUDE MYTHOS OUTPUT > > One of the first confirmed public outputs from Mythos. It's pretty insane. Just with a simple prompt. Better than Gemini with SVG's. Google is cooked. > > A lot more coming soon. [pic.twitter.com/nq0KPsIXwN](https://t.co/nq0KPsIXwN?ref=testingcatalog.com) > > — can (@marmaduke091) [June 5, 2026](https://x.com/marmaduke091/status/2062799040689893500?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Red-team access has typically preceded a wider rollout by a week or two, suggesting a plausible launch in the second half of June, close to when OpenAI's rumored GPT-5.6 (codename kindle-alpha) is also expected. > 🚨 NEW GPT-5.6 CHECKPOINTS DROPPED > > OpenAI is testing 2 new checkpoints: > > \> - kindle-alpha (release candidate) > \> - kepler-alpha > > Join Our server we are serving back to back test before launch 🫣 [pic.twitter.com/oalhzIV092](https://t.co/oalhzIV092?ref=testingcatalog.com) > > — Chetaslua (@chetaslua) [June 5, 2026](https://x.com/chetaslua/status/2062805007779651886?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Early discussions, including creative venues like VoxelBench, suggest that Oceanus outputs are significantly superior to those of current models, even with minimal effort, though nothing has been independently confirmed. The leap looks promising, though not yet settled. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my1.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my2.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my3.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my4.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my6.webp) ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/my8.webp) Oceanus on [VoxelBench](https://x.com/voxelbench?ref=testingcatalog.com) Earlier reports initially linked [Mythos to Claude Code and Claude Security](https://www.testingcatalog.com/anthropic-prepares-mythos-1-for-claude-code-and-claude-security/), targeting developers and security teams. Whether it will also reach business, personal, or Max tiers, or remain enterprise-only, is unclear; Pro almost certainly waits. [![CTA Image](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/e4995a1dd8fdcdb222d3aa03c7303bc8.jpg)](https://discord.gg/DevMode?ref=testingcatalog.com) Join Dev Mode Discord for more! [Join ](https://discord.gg/DevMode?ref=testingcatalog.com) This development coincides with Anthropic's new Institute paper, which argues that AI is already accelerating AI development, citing Mythos Preview achieving a 52x training-optimization speedup. Notably, the company frames this as a cautionary note, emphasizing that the field has not yet achieved recursive self-improvement and calling for verifiable methods to slow down, rather than celebrating prematurely. ### OpenSquilla lets AI agents organize their own skills URL: https://www.testingcatalog.com/opensquilla-lets-ai-agents-organize-their-own-skills/ Last updated: 2026-06-05T11:25:57.000Z Most agent projects are still racing to field a smarter chat loop or a longer list of tools. OpenSquilla, an Apache-licensed, self-hostable runtime, is making a quieter bet: that the next round of efficiency comes from the harness rather than the model. Its founding idea was cost-aware routing, scoring each turn and sending trivial work to cheap models while reserving heavier reasoning for tasks that warrant it. That is becoming table stakes. The part worth watching now is MetaSkill, the project's attempt to let an agent organize its own capabilities rather than rely on hand-written workflows. ![OpenSquilla](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/screenshot-4x-viewport-2026-06-05T10-07-37.webp) The premise is the combinatorial problem facing every maturing agent: writing a single-task skill is trivial, but composing hundreds of community skills into something reliable collapses into guesswork once real complexity arrives. MetaSkill answers with a meta-protocol, a markdown spec that tells the model how to discover, rank, and compose atomic skills, declaring the resulting workflow in a structured header that the runtime validates before anything runs. A goal described in plain language becomes an inspectable, replayable execution chain rather than a one-off prompt. > Tonight, as promised 🦞 > > That volcano plan this morning? Not a chatbot — it was > MetaSkill, OpenSquilla's self-organizing skill protocol. > > You describe the goal in plain words. It discovers, picks, > and composes the right skills into a real, safe workflow — > and it can even write… [pic.twitter.com/f1ksWydMp5](https://t.co/f1ksWydMp5?ref=testingcatalog.com) > > — OpenSquilla (@OpenSquilla) [June 1, 2026](https://x.com/OpenSquilla/status/2061481500974157956?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The runtime ships with ready-made workflows for jobs such as research-to-report and project planning, and during idle time, it revisits its execution traces, distills recurring patterns, and drafts candidate workflows. The catalog grows in the background. That is also where the open questions sit. The headline savings figures are the project's own benchmarks, not independent results, and machine-composed workflows raise obvious reliability and safety concerns that the design seeks to contain through proposal gates, tool allowlists, and syscall-level sandboxing. SPONSORED Star the repo & get OpenSquilla [Learn more ](https://github.com/opensquilla/opensquilla?ref=testingcatalog.com) Strategically, [OpenSquilla](https://www.testingcatalog.com/opensquilla-launches-open-source-ai-agent-to-cut-token-costs/) is positioning against heavier agents such as OpenClaw, even shipping migration tooling, while aligning with the wider move toward portable skill specifications. For teams running long-horizon agents where token bills compound, the pitch lands squarely at the orchestration layer. Routing, tiered memory, and the first MetaSkill capabilities sit in the public releases today; a further iteration looks set to follow from active development. If orchestration rather than model size proves the durable lever, projects like this reframe where the real moat lies. ### Microsoft Build 2026 recap, from Windows to Copilot, all AI URL: https://www.testingcatalog.com/microsoft-build-2026-recap-from-windows-to-copilot-all-ai/ Last updated: 2026-06-03T20:18:17.000Z Microsoft used its Build 2026 keynote to make its clearest case yet for owning the models beneath its products, rolling out seven new MAI systems across reasoning, coding, image, voice, and transcription. The headline is MAI-Thinking-1, a mid-sized 35-billion-parameter reasoning model with a 256K context window that Microsoft says was built without distillation. > Link [https://t.co/VFu7G61PFH](https://t.co/VFu7G61PFH?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 2, 2026](https://x.com/testingcatalog/status/2061848629627809821?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The company claims blind raters prefer it to Sonnet 4.6 and that it matches Opus 4.6 on SWE-Bench Pro, though it targets enterprise and sits in private preview on Foundry behind an access request. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-20.22.36.webp) New MAI models Alongside it, MAI-Code-1-Flash started rolling out in VS Code through the GitHub Copilot model picker, tuned for fast, low-cost coding and pitched above Claude Haiku 4.5 on price-to-performance. Neither is a frontier model, and both would have landed as stronger debuts a year ago, yet they read as a real first step: Microsoft now wants to power its own products with its own models and train them on Microsoft-specific use cases rather than leaning on OpenAI. Both are worth testing once they settle, and the reasoning model in particular will be worth watching over time. > MICROSOFT 🔥: New MAI Code 1 Flash and MAI Thinking 1 models have been revealed on the official MAI website! > > Also, MAI Image 2.5, MAI Voice 2, and MAI Transcribe 1.5 are there too. > > \> MAI-Code-1-Flash plans and reasons through complex coding tasks from start to finish, so you… [pic.twitter.com/oqhjWoP0JY](https://t.co/oqhjWoP0JY?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 2, 2026](https://x.com/testingcatalog/status/2061869964840341787?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The most competitive release is MAI-Image-2.5 and its Flash variant, which has sat on LM Arena for a while and, to my eye, lands on par with or in places ahead of Nano Banana Pro. It is already live in PowerPoint and reaching OneDrive, and it is certainly worth a try. The voice model still needs more hands-on time, but it matters because it underpins a wave of voice-first hardware. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-19.15.57.webp) Hello for Business AI assistant ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-19.16.49.webp) AI Badge device Microsoft showed a home display that sits beside you and runs your agents by voice, likely answering back the same way and grounded in WorkIQ data, plus an AI PC with a camera and voice control for driving agents. For anyone living inside the Microsoft ecosystem, those are worth watching. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-18.43.28.webp) Agentic Windows Terminal Around them sat the heavier infrastructure story: the NVIDIA collaboration, new Surface laptops and chips, and fresh progress in quantum computing. The other anchor was the [Copilot](https://www.testingcatalog.com/tag/microsoft-copilot/) super app, expected to fold chat, Cowork, and GitHub-based coding into a single shell, with long-running "autopilots" layered on top. Scout is the first of those, an always-on agent powered by OpenClaw that runs across Teams, Outlook, and the desktop with a governed Entra identity per agent. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-18.37.39.webp) New Aion models Developers also got new local on-device models for Windows and a native agentic terminal powered by GitHub Copilot, which should land well given how close it sits to the operating system. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-19.57.00.webp) Github Copilot app OpenClaw itself is coming natively to Windows with a dedicated app and a sidebar widget that can be toggled on or off and surface status as it works through tasks, with plenty of configuration inside its chat. Microsoft even demoed its guardrails by walling OpenClaw off from certain folders to see whether it would try to delete files it could not touch, and the controls held up well. Founder Peter Steinberger taking the stage was a nice moment in the run of the show. ![Microsoft Build 2026 recap ](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-at-19.41.02.webp) OpenClaw on Windows As a whole, it was a strong outing and an important step for Microsoft, if a touch long. Set against Google I/O, where updates were packed into tight segments, Microsoft stretched things out and saved its most interesting reveals for the end, which made the opening stretch a harder watch for anyone not deep in enterprise. There is plenty here worth getting hands on, OpenClaw and the super app, most of all, so the wait now is for access. ### OpenAI makes its next hardware move with Opal Electronics URL: https://www.testingcatalog.com/openai-makes-its-next-hardware-move-with-opal-electronics/ Last updated: 2026-06-03T19:32:35.000Z OpenAI is investing fresh capital into hardware by leading a new funding round for Opal Electronics, a San Francisco startup renowned for its high-end webcams. Known for creating the C1 and the pocket-sized Tadpole, Opal is reportedly preparing to launch a new product line in the coming months. This move will extend Opal's reach beyond cameras and into AI-native devices designed for creative work. > the table. > an update on opal electronics. [https://t.co/91eGxO5qV6](https://t.co/91eGxO5qV6?ref=testingcatalog.com) > > — Opal (@opalelectronics) [June 2, 2026](https://x.com/opalelectronics/status/2061836955659485312?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The specifics of this new device remain unconfirmed. However, Opal’s background suggests it will be vision- and capture-oriented, potentially utilizing OpenAI’s image, video, and real-time voice models as its foundation. Integrating voice into a physical product would provide OpenAI with valuable insights into how users interact with an always-listening companion, offering data that a chat window cannot supply. Details such as form factor, pricing, and exact capabilities are still under wraps. This investment aligns with a broader trend. Over the past year, OpenAI has been focusing on hardware development, inspired by Sam Altman’s vision for “ambient computing.” This concept involves lightweight gadgets that can sense the world in real time without relying on a screen. OpenAI's most notable project in this area, a palm-sized, screenless device developed with Jony Ive following the multibillion-dollar io acquisition, has faced delays. Originally slated for release, it has now been pushed to 2027 due to software, privacy, and computational challenges, and has lost its original name due to a trademark dispute. Despite these setbacks, Chris Lehane has maintained that devices remain a top priority for the company through 2026. > 👀 Interesting... Google AI Studio product lead quoting [https://t.co/AKoFUgA9O6](https://t.co/AKoFUgA9O6?ref=testingcatalog.com) website, who just announced funding from [@openai](https://x.com/OpenAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [https://t.co/0gIK48vVJf](https://t.co/0gIK48vVJf?ref=testingcatalog.com) [pic.twitter.com/3gsbweqon4](https://t.co/3gsbweqon4?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [June 3, 2026](https://x.com/testingcatalog/status/2062251883914272941?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) In this context, the partnership with Opal appears to be more than a standalone venture; it is a strategic move to explore form factors, expedite product release, and gather user feedback while the flagship project continues to develop. For a company intent on moving beyond the smartphone, investing in the hardware that will eventually replace it is a logical progression, making this development one to watch as the first products begin to emerge. [Source](https://www.wired.com/story/opal-electronics-openai-investment-ai-powered-audio-gadget/?ref=testingcatalog.com) ### Google tests Planning Mode for NotebookLM Video Overviews URL: https://www.testingcatalog.com/google-tests-planning-mode-for-notebooklm-video-overviews/ Last updated: 2026-06-02T22:07:16.000Z Google looks set to bring a planning mode to NotebookLM’s Video Overviews, a control that has surfaced inside a recent build of the research tool. The option lives in the customization menu, the same panel, reached via the pencil icon on the Video Overview tile, where users already pick a format, choose a visual style, and write a custom prompt. A new toggle would switch on the planning step. The exact layout is still taking shape, but the behavior closely resembles the plan-then-build pattern familiar from coding assistants. Rather than sending a prompt straight to rendering, Gemini would first draft a plan for what the video should cover and pause to let the user edit and approve it before generation begins. For educators, researchers, students, and anyone turning dense sources into watchable explainers, that checkpoint matters: it adds editorial oversight over structure and pacing, and it heads off wasted generations on a clip that misses the point. Today, [NotebookLM](https://www.testingcatalog.com/tag/notebooklm/) hands those calls to Gemini, acting as a silent creative director, with no place for the user to step in. The toggle also hints at something larger under the hood. The capability aligns with [Gemini Omni](https://www.testingcatalog.com/google-rolls-out-gemini-omni-ai-for-video-generation-and-editing/), the multimodal model Google introduced at I/O 2026 in May, which now serves as its default video engine and can generate explainer-style clips from a single prompt. Moving Video Overviews from the current Veo-based stack to Omni would fit Google’s push to consolidate text, image, and video into one “anything from anything” system, and a planning step is the kind of control Omni’s editing-first design naturally supports. Such a shift would likely accompany any broader upgrade to the video pipeline. No timeline has been attached to the feature, and the planning mode remains in development for now. Tip by [@thomas\_gmry](https://x.com/thomas%5Fgmry?ref=testingcatalog.com) ### TinyFish Bigset turns text prompts into live datasets from web URL: https://www.testingcatalog.com/tinyfish-bigset-turns-text-prompts-into-live-datasets-from-web/ Last updated: 2026-06-02T17:36:46.000Z [TinyFish](https://www.tinyfish.ai/?ref=testingcatalog.com) has launched [Bigset](https://github.com/tinyfish-io/bigset?ref=testingcatalog.com), an open-source multi-agent system that turns a plain-language sentence into a structured dataset pulled from the live web. You describe what you want, and Bigset infers the schema, sends autonomous agents to research it on real web pages, verifies their findings against sources, deduplicates, and hands back a clean table you can export as CSV or XLSX. Set a refresh cadence from 30 minutes to weekly, and have the agents rerun on schedule so the dataset stays current without anyone needing to touch a script. 0:00 /0:48 1× University Application Tracker The work is split across two agent roles. An orchestrator agent does breadth-first discovery, identifying which rows belong in the dataset and where on the web to find them, then dispatches sub-agents to fill each one. The orchestrator holds no write access of its own. Each sub-agent researches a single entity under a tight budget of 6 tool calls, pulls real data via TinyFish Search and Fetch, and inserts one verified row with its source URLs and a record of how the data was found. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-113120.png) TinyFish Bigset Sub-agents are instructed never to fabricate values, to leave fields blank when they cannot be confirmed, and to reject duplicate primary keys automatically. The orchestrator runs until the dataset reaches its row target, building faster as it learns where the data lives. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/06/Screenshot-2026-06-02-113036.png) Tinyfish Bigset Results Bigset is licensed under AGPL-3.0 and runs self-hosted through Docker, with schema inference on Claude Sonnet 4.6 and the agent roles on Qwen3.7-max by default, all routed through OpenRouter and configurable per role. The team is candid that the project is experimental: a dataset takes 2 to 5 minutes to build, it works best on topics with public web data, and the free tier covers 2,500 row operations per month. It ships with 9 curated public datasets covering AI companies hiring, GPU prices, model pricing, and top open-source repositories, browsable without an account. ****SPONSORED** Test it out for yourself on TinyFish! [Take me there! ](https://bit.ly/4x6SIyk?ref=testingcatalog.com) TinyFish is the Palo Alto-based company behind the platform, backed by $47 million in Series A funding led by ICONIQ, and counts Google, DoorDash, and Amazon among its enterprise clients, having processed more than 40 million agent operations. Bigset is built directly on TinyFish Search and Fetch, the same web infrastructure underneath the company's enterprise agent products, and arrives as the open-source answer to proprietary natural-language dataset tools, with no per-seat pricing, no domain restrictions, and full pipeline ownership for anyone who runs it themselves. Star it on [Github](https://github.com/tinyfish-io/bigset?ref=testingcatalog.com) and grab an [API key](https://bit.ly/4x6SIyk?ref=testingcatalog.com)! 🔥 ### What releases to expect from Anthropic in coming weeks URL: https://www.testingcatalog.com/what-releases-to-expect-from-anthropic-in-coming-weeks/ Last updated: 2026-05-31T13:48:26.000Z Fresh off Opus 4.8, which landed roughly a week after Google I/O, Anthropic finds itself in a very strong position. The model arrived only about six weeks after Opus 4.7, posting category-leading scores on agentic coding and reasoning while shipping effort controls and a faster mode alongside it. Pair that cadence with a funding round near a trillion dollars, one that lifted the company past OpenAI for the first time, and with [**Mythos-grade models**](https://www.testingcatalog.com/anthropic-prepares-mythos-1-for-claude-code-and-claude-security/) expected within weeks, and the backdrop for what comes next looks formidable. ![Mythos 1](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Claude-05-24-2026_12_04_AM.jpg) Beneath the model news sits a deeper story: a cluster of products, surfaced through code references and hidden interface strings, that push Claude well past the chat window. The centerpiece is [**Conway**](https://www.testingcatalog.com/anthropics-works-on-its-always-on-agent-with-new-ui-extensions/), an always-on agent that runs inside a managed container and closely mirrors the workbench setup. It would appear as a separate sidebar option opening a dedicated page tied to a "Conway instance," a standalone environment rather than a chat view. Inside, users could connect integrations, install skills, and add plugins, then arrange them as switchable tiles or tabs across a side panel, moving between tools as though tabbing through a browser. ![Conway](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Conway-Claude-04-02-2026_01_04_AM.jpg) Conway A custom abstraction labeled "UI tabs" points to a new extension standard, possibly a .EXT package format, that Anthropic may open up so people can build, share, and download their own or third-party add-ons from marketplaces. Those extensions could prove reusable across any agent built atop Claude-managed agents. Webhook support hints at public URLs that wake the instance when outside services call, alongside Chrome control and notifications. Conway looks bound for Claude Code and mobile, with web likely given its remote footing, and would cap each person at a single agent, a direct answer to OpenClaw and Hermes, whose founders and momentum have drawn plenty of attention this year. ![File Memory](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Screenshot-2026-05-23-at-22.13.10--1--1.png) File Memory [**Memory**](https://www.testingcatalog.com/anthropic-plans-claude-memory-update-with-new-memory-files/) is shifting in parallel. A file-based option would let managed agents structure, categorize, and refine stored context over time rather than writing a flat summary, with files optimized as they accumulate. The payoff would be a shared memory layer running across products. ![Orbit](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Claude-05-05-2026_01_42_AM--1--1.jpg) Orbit [**Orbit**](https://www.testingcatalog.com/anthropic-is-working-on-orbit-its-upcoming-proactive-assistant/), an upcoming proactive assistant, would allow Claude to capture personalized insights from different sources and proactively reach out to the user when needed. > Users will get personalized insights from Gmail, Slack, GitHub, Calendar, Drive, Figma, and other apps, which Claude will generate proactively. According to previously discovered strings, Claude Orbit will gain the ability to "Deploy favorite apps." ![Operon](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Claude-03-27-2026_11_59_PM.jpg) Operon [**Operon**](https://www.testingcatalog.com/anthropic-tests-claude-operon-for-scientific-research-in-biology/), meanwhile, targets life sciences researchers: a fourth desktop mode alongside Chat, Code, and Cowork, offering a private environment, project sessions, Plan and Auto modes, and local file access for work such as CRISPR screen design or single-cell RNA analysis. Given Anthropic's prior AI for Science and Life Sciences efforts, piloting with select organizations already looks plausible. ![BugCrawl](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Claude-04-24-2026_12_42_AM.jpg) BugCrawl [**BugCrawl**](https://www.testingcatalog.com/anthropic-tests-new-bugcrawl-tool-for-claude-code-bug-detection/) rounds out the agentic push, mirroring Claude Security but hunting general bugs rather than vulnerabilities. It surfaces as its own Claude Code entry with a repository picker and a high token-burn warning, and would likely pull tickets from GitHub, Jira, or Linear before adding tests, verifying the fix, and watching the rollout. Several smaller threads complete the picture: 1. A refreshed [**Claude Security dashboard**](https://www.testingcatalog.com/anthropic-prepares-mythos-1-for-claude-code-and-claude-security/) appears due. 2. Meeting-note capture in the vein of Granola or Notion. 3. [**AI Fluency**](https://www.testingcatalog.com/anthropic-to-introduce-personal-ai-fluency-scorecard-in-claude/) scorecard in settings, along with other stats. 4. [**A new voice mode**](https://www.testingcatalog.com/anthropic-plans-expanding-claude-voice-mode-to-more-languages/) covering many more languages, still text-to-speech, with one or two voices each, yet able to switch tongues mid-sentence. Current tests reportedly lean on Haiku 4.5, which has unsettled some, though internal wiring can change before anything ships. Besides that, pixelated avatars spotted earlier seem to have been pulled and should not be expected soon. ### Exclusive: New screenshots of upcoming Copilot Super App URL: https://www.testingcatalog.com/exclusive-new-screenshots-of-upcoming-copilot-super-app/ Last updated: 2026-06-23T23:15:41.000Z UPDATE: Copilot super app, along with a new Scout agent have been officially announced during Microsoft Build 2026! > Meet Microsoft Scout. > > An always-on agent that keeps work moving, taking action without needing to be prompted each time. > > As Microsoft’s first Autopilot agent, Microsoft Scout works across Teams, Outlook, OneDrive, and more—taking action within the controls your organization… [pic.twitter.com/YqeDABRHAy](https://t.co/YqeDABRHAy?ref=testingcatalog.com) > > — Microsoft 365 (@Microsoft365) [June 2, 2026](https://x.com/Microsoft365/status/2061874930547871868?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) ### The Story Microsoft looks set to use its Build conference on June 2 in San Francisco to lay the groundwork for the unified [Copilot super app](https://fortune.com/2026/05/29/microsoft-working-on-super-app/?ref=testingcatalog.com) it has been building under the internal slogan “Delivering one Copilot.” > Loud and clear. [#MSBuild](https://x.com/hashtag/MSBuild?src=hash&ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) kicks off on June 2\. [pic.twitter.com/QU9cAbG3Lr](https://t.co/QU9cAbG3Lr?ref=testingcatalog.com) > > — Microsoft (@Microsoft) [May 29, 2026](https://x.com/Microsoft/status/2060398263665078476?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Freshly surfaced screenshots provide the clearest look yet at the shell. Earlier views showed the Scout always-on agent within an **Autopilot** section; **two new tabs now complete the picture.** > Sources: This is a leaked screenshot of Microsoft's coming super app redesign of Copilot, featuring an OpenClaw-like agent called Scout that it's planning to announce soon [https://t.co/4IvkdOhIf8](https://t.co/4IvkdOhIf8?ref=testingcatalog.com) [pic.twitter.com/vew3fQCUAM](https://t.co/vew3fQCUAM?ref=testingcatalog.com) > > — Alex Heath (@alexeheath) [May 30, 2026](https://x.com/alexeheath/status/2060516372891992097?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The first is a coding surface with the **GitHub Copilot** mark, closely mapping to the Claude Code panel in the Claude app. It allows users to pick a work tree, points to both remote environments and local repositories, carries a model selector, lists every repo, and adds Routines, a scheduled-task layer built for code. Sitting atop GitHub Copilot and its millions of paying developers, it could be a real upgrade for teams already standardized on GitHub, especially once [Microsoft’s own coding model](https://www.theinformation.com/newsletters/ai-agenda/microsoft-release-new-coding-model-next-week-comeback-attempt?utm%5Fcampaign=Brand+Partnerships&utm%5Fcontent=PwC&utm%5Fmedium=organic%5Fsocial&utm%5Fsource=bluesky%2Cfacebook%2Cthreads%2Ctwitter&rc=hnvqnq) arrives tuned for GitHub Copilot-specific tool use. ICYMI: Microsoft is preparing [upgrades for image and voice models](https://www.testingcatalog.com/microsoft-readies-new-mai-voice-and-image-models-for-build-2026/), too. ![Exclusive: Copilot Cowork tab of Copilot super app](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/IMG_7088.JPG) Exclusive: Copilot Cowork tab of Copilot super app The second tab, **Cowork**, pulls from several sources, aggregates the data, and proposes prompts like preparing for the week from a calendar or researching a company, similar to Copilot’s current document and presentation work. The open question is about local files: the screenshot shows it running in Edge via a URL, so whether it reaches the desktop or remains fully remote is unconfirmed. A sidebar with Library and Projects keeps these jobs apart from plain chat, coding, and Autopilot. The division matters because the company, with Jacob Andreou now leading [Copilot](https://www.testingcatalog.com/tag/microsoft-copilot/) after a reshuffle, is trying to boost weak adoption by folding scattered tools into a single home, just as OpenAI and Anthropic converge on the same always-on, multi-mode pattern. Teams integration hints that Scout could run remotely, the nearest Microsoft may get to the chat-app control that made such agents popular. A nod at Build looks probable, though the app itself is aimed at late summer. ### Microsoft released new MAI voice and image models for Build 2026 URL: https://www.testingcatalog.com/microsoft-readies-new-mai-voice-and-image-models-for-build-2026/ Last updated: 2026-06-02T21:29:46.000Z UPDATE: Microsoft has announced 7 new AI models during Microsoft Build 2026 - MAI Image 2.5, MAI Image 2.5 Flash, MAI Voice 2, MAI Voice 2 Flash, MAI Transcribe 1.5, MAI Code 1 Flash, and MAI Thinking 1. > Seven new models launching at Build: let’s go! > Reasoning. Code. Image. Transcribe. Voice. > > Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models > > Thread 🧵 [#MSBuild](https://x.com/hashtag/MSBuild?src=hash&ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [pic.twitter.com/g3WQIcIQ24](https://t.co/g3WQIcIQ24?ref=testingcatalog.com) > > — Microsoft AI (@MicrosoftAI) [June 2, 2026](https://x.com/MicrosoftAI/status/2061887500541366489?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) ### The Story Microsoft heads into its Build conference on June 2 in San Francisco with more in its model pipeline than the [MAI-Image-2.5](https://microsoft.ai/news/mai-image-2-5-launches-at-no-3-on-arena-ai/?ref=testingcatalog.com) that it has already shown on Arena, where the text-to-image system landed third behind OpenAI’s gpt-image-2 and Google’s Nano Banana 2\. That release is lined up for the MAI Playground and Foundry, but three additional models are taking shape within the company’s stack, none of which are publicly available yet. > Exciting news, MAI-Image-2.5 (Preview) from [@MicrosoftAI](https://x.com/MicrosoftAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) debuts at #3 in the Text-to-Image Arena with a score of 1,254 — a +72 point improvement over MAI-Image-2. > > A top 5 arena previously held only by [@GoogleDeepMind](https://x.com/GoogleDeepMind?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) and [@OpenAI](https://x.com/OpenAI?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) has a new lab in the mix. > > Congrats to the… [https://t.co/stHydZYbNN](https://t.co/stHydZYbNN?ref=testingcatalog.com) [pic.twitter.com/4eVXxfbI6M](https://t.co/4eVXxfbI6M?ref=testingcatalog.com) > > — Arena.ai (@arena) [May 26, 2026](https://x.com/arena/status/2059346024632820146?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The first, **MAI-Transcribe-1.5**, is a modest step up from the speech-to-text model launched in April, which already claimed the lowest word error rate across 25 languages. The image side draws more attention: **MAI-Image-2.5** looks set to ship in two variants, a high-quality version and a faster one labeled **MAI-Image-2.5e**, mirroring the split seen with MAI-Image-2\. It would also accept image uploads, opening the model to editing as well as generation, putting it on par with rivals from Google and OpenAI. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/MAI-Playground-Microsoft-AI-05-30-2026_10_20_PM.jpg) The most striking find is **MAI-Voice-2**, a multilingual successor to the company’s text-to-speech model. While MAI-Voice-1 began in English, the new version adds German, Australian and US English, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Dutch, Portuguese, Turkish, Vietnamese, and Chinese, with a wider emotional range that covers tones such as angry, confused, and embarrassed. Early samples suggest it can whisper, too. Harper whisper egret 0:00 /18.8 1× Ethan shouting egret 0:00 /15.696 1× Field isla joyful 0:00 /16.176 1× All three would feed Copilot, Teams, and Azure Speech, and fit the developer crowd that Build is made for. The timing matches a broader push, as Mustafa Suleyman’s team weans the company off OpenAI following April’s renegotiation. Reports point to a homegrown [coding model](https://www.theinformation.com/newsletters/ai-agenda/microsoft-release-new-coding-model-next-week-comeback-attempt?utm%5Fcampaign=Brand+Partnerships&utm%5Fcontent=PwC&utm%5Fmedium=organic%5Fsocial&utm%5Fsource=bluesky%2Cfacebook%2Cthreads%2Ctwitter&rc=hnvqnq) for GitHub Copilot at the show, too, while a Copilot “[super app](https://sources.news/p/leaked-microsoft-ai-copilot-super-app-autopilot-scout?ref=testingcatalog.com)” that integrates chat, coding, and agents into a single hub is expected later in the summer. ### 3 upcoming NotebookLM features we all should be waiting for URL: https://www.testingcatalog.com/3-upcoming-notebooklm-features-we-all-should-be-waiting-for/ Last updated: 2026-05-30T19:23:36.000Z Google appears to be lining up a batch of NotebookLM features that have been in the works for months, surfacing quietly in recent builds even as the team drops hints that an announcement may not be far off. Three additions stand out, and together they sketch a clear direction. > as is tradition, NotebookLM’s big update is not on IO day… > > Team has cooked. > > — Simon (@tokumin) [May 24, 2026](https://x.com/tokumin/status/2058348822993174666?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The first is [**Personal Preferences**](https://www.testingcatalog.com/google-tests-personal-intelligence-for-notebooklm-conversations/), which debuted in Gemini and is now set to reach NotebookLM. It would let the tool learn from your activity and build editable personas, adjusting tone and technical depth to how you work. While Gemini’s version reaches into Gmail, Drive, Photos, and Calendar, the NotebookLM signals so far lean toward in-app personalization drawn from your notebooks and chats, which is useful for anyone doing repeated, deep-context research. > Allow NotebookLM to use your past interactions (e.g., conversations, artifacts, and customization instructions) to understand your preferences and tailor the experience to your needs. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/NotebookLM-05-30-2026_12_53_AM.jpg) [**Connectors**](https://www.testingcatalog.com/google-tests-canvas-and-connectors-on-notebooklm/), sitting alongside it in settings, would close that gap. Working much like MCP, it would pull outside data into a notebook, most likely starting with Google’s own services such as Calendar, Gmail, and Drive. The piece is not yet operational, and the roster of supported sources remains open. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/NotebookLM-05-30-2026_12_53_AM--1-.jpg) [**Canvas**](https://www.testingcatalog.com/google-tests-canvas-and-connectors-on-notebooklm/) is arguably the headline. Found in the Studio panel, it would turn sources into a custom artifact — an interactive timeline, an explainer web page, a lightweight game, or a visualizer, guided by a prompt describing what you want and how. It extends the outputs that NotebookLM already offers, including infographics, slide decks, data tables, and mind maps. ![NotebookLM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/NotebookLM-05-30-2026_01_00_AM.jpg) The combination is where it gets compelling. With NotebookLM now living inside Gemini, the three would let people work across their sources without copying material between tools, matching Google’s push to turn a source-grounded reader into a workspace for building structured, visual experiences on top of documents. On models, [NotebookLM](https://www.testingcatalog.com/tag/notebooklm/) moved to Gemini 3 late last year; with Gemini 3.5 Flash now the global default after I/O 2026, the Flash branch of that family is the natural next base. A firm timeline is still missing, so the open question is when, not whether. ### Perplexity tests daily Digest feature for its Computer agent URL: https://www.testingcatalog.com/perplexity-tests-daily-digest-feature-for-its-computer-agent/ Last updated: 2026-05-30T13:01:10.000Z Perplexity appears to be preparing a [digest](https://www.testingcatalog.com/perplexity-prepares-digest-feature-for-personalized-summaries/) experience for its Computer agent, with recent builds showing a redesigned page that is currently empty but is wired to populate with daily updates pulled from a range of sources. The shell is in place ahead of the content, which suggests the feature is still being assembled rather than nearing a public switch-on. Alongside the page is a fresh settings tab dedicated to the digest, letting users shape what they see and govern which connectors feed it. The roster spans everyday office tooling, such as email and cloud drives, through to developer-leaning services like Linear, GitHub, and Notion, a mix suggesting Perplexity wants the briefing to draw from both inbox chatter and active project work. ![](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Perplexity-Digest-05-29-2026_04_22_PM-1.jpg) Memory looks set to play a central role. The company has been reworking its recall layer toward a knowledge-based structure reminiscent of the file-driven, categorized stores that always-on agents like OpenClaw and Hermes rely on, where separate documents hold distinct slices of context. Insights surfaced from that memory would be folded into each digest, giving the summary a personal cast rather than a generic news roll. One capability tagged “coming soon” is the option to route the digest to Slack, an obvious draw for teams and anyone leaning on Perplexity at work, and a natural extension of Computer’s existing Slack presence, which already answers DMs and runs scheduled workflows. No firm date has emerged, though the work has reached production testing, which usually means word lands within weeks. Because the digest leans on [Perplexity](https://www.testingcatalog.com/tag/perplexity/) Computer and its credit-metered compute, the likeliest home is the Max tier; Pro access feels less certain, given Computer’s wider rollout, but that remains speculation for now. Taken together, the moves push Computer further toward an ambient briefing layer that watches your tools, remembers your context, and reports back on a schedule. ### Explee launches AutoGTM AI Agent for outbound sales URL: https://www.testingcatalog.com/explee-launches-autogtm-ai-agent-for-outbound-sales/ Last updated: 2026-05-29T11:06:37.000Z Explee has made AutoGTM publicly available, a 24/7 AI sales agent that converts a website URL into a running outbound pipeline without any manual prospecting, template setup, or SDR coordination. The pitch is direct: paste your domain, and a team of seven autonomous AI agents researches your market, identifies ideal customer profiles, locates matching companies, verifies contact data, writes personalized cold emails, and sends them on a schedule. The process from URL input to first outreach completes in roughly two minutes. ![AutoGTM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/AutoGTM-by-Explee-----24-7-AI-agent-that-finds-clients-while-you-sleep-05-29-2026_12_30_PM.jpg) The system runs against a database of over 105 million company profiles and 536 million people profiles, with verified email addresses and pre-warmed mailboxes ready for outreach from day one. AutoGTM reports a 97% email deliverability rate across its automated sequences, which include follow-ups handled without human involvement. The agent layer conducts in-depth research to personalize each message based on the prospect's company and contact, rather than inserting a name into a shared template. Calendar and CRM integration is included alongside API access for teams building AutoGTM into a broader agent stack. 0:00 /0:21 1× Pricing operates on a pay-as-you-go model at $0.03 per email sent, with daily budget caps and the option to scale spend without a plan change. Explee positions this at roughly 15 times below the cost of traditional data providers like ZoomInfo and Apollo, which charge flat subscription rates for data that still requires a separate outbound motion layered on top. New accounts receive $50 in free credits to start. The product is available globally online with no minimum commitment. ![AutoGTM](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/AutoGTM-by-Explee-----24-7-AI-agent-that-finds-clients-while-you-sleep-05-29-2026_12_31_PM.jpg) AI-native founders and solo product teams frequently reach a point where product development runs entirely on AI tooling, while outbound sales remain manual. Sourcing leads, validating emails, writing cold messages, and managing follow-up sequences are time-consuming tasks that do not compound. AutoGTM covers the full sales development loop from the moment a company URL is entered, with the intended audience being founders, AI-native builders, and small teams that want a qualified pipeline without having to build or staff a sales function. ****SPONSORED** Test it out for yourself on Explee! [Take me there! ](https://explee.com/auto-gtm/x/an1?ref=testingcatalog.com) Explee is a London-based company incorporated as Explee LTD, operating from International House on Essex Street. The company was founded as a semantic search engine for B2B companies and decision-makers, with a proprietary database that has grown to include over 105 million company profiles globally. AutoGTM extends that infrastructure into a fully autonomous outbound motion, collapsing what was previously a stack of separate tools for search, enrichment, email tooling, and sequencing into a single agent loop that operates continuously. ### Anthropic launches dynamic workflows for Claude Code URL: https://www.testingcatalog.com/anthropic-launches-dynamic-workflows-for-claude-code/ Last updated: 2026-05-29T10:23:18.000Z Anthropic has announced the introduction of dynamic workflows within Claude Code, a feature designed to handle complex engineering tasks at scale. This new capability allows users to orchestrate large, multi-step coding projects through parallelized agents that operate on subtasks simultaneously. For example, the porting of Bun from Zig to Rust, involving around 750,000 lines and rigorous test suite validation, was accomplished by leveraging these workflows. > Also new in Claude Code: dynamic workflows (research preview). > > For the hardest tasks, Claude makes a plan, runs hundreds of parallel subagents, and verifies its work before reporting back. Think a migration touching hundreds of files. > > Read more: [https://t.co/7gt06kGkDN](https://t.co/7gt06kGkDN?ref=testingcatalog.com) > > — Claude (@claudeai) [May 28, 2026](https://x.com/claudeai/status/2060042710753382816?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The system automatically plans and distributes work, checks results for accuracy, and iterates until consensus is reached, ensuring coordination even for jobs that span several days. Users are prompted for confirmation before a workflow executes, and organization admins can manage access and settings. > Excited to share our most powerful new Claude Code feature: dynamic workflows! > > Mention "workflow" in a prompt and Claude will dynamically create an orchestration plan that it strictly follows, allowing you to confidently trust that every stage happens in the right order even… [pic.twitter.com/ilvthUbr82](https://t.co/ilvthUbr82?ref=testingcatalog.com) > > — cat (@\_catwu) [May 28, 2026](https://x.com/%5Fcatwu/status/2060054180379689074?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [Anthropic's](https://www.testingcatalog.com/tag/claude/) Claude Code platform, known for its AI-powered coding assistance, is now targeting power users, including engineering teams and organizations managing large-scale projects. Dynamic workflows are immediately available to those on Max and Team plans or via the API, with Enterprise customers able to opt in by enabling the feature through admin controls. This launch positions Anthropic competitively against other AI coding assistants by scaling to more demanding workflows and offering features such as workflow recovery and granular administrative oversight. Early user feedback highlights the tool's ability to expedite previously time-intensive engineering processes. [Source](https://claude.com/blog/introducing-dynamic-workflows-in-claude-code?ref=testingcatalog.com) ### Nano Banana 2 and Nano Banana Pro are now in General Availability URL: https://www.testingcatalog.com/nano-banana-2-and-nano-banana-pro-are-now-in-general-availability/ Last updated: 2026-05-29T10:12:20.000Z Google Cloud has launched Nano Banana 2 (Gemini 3.1 Flash Image) and Nano Banana Pro (Gemini 3 Pro Image), now generally available through the Gemini Enterprise Agent Platform. This release directly targets enterprise customers seeking robust, secure, and scalable image generation and editing capabilities. Both models are built on Google Cloud’s infrastructure, ensuring reliability and compliance for businesses operating at scale. The models support 1K and 2K image output, with a 4K option in preview, and offer expanded functionality including a new preview feature allowing video files to be used as input prompts for context-aware image creation. This positions the models as versatile tools for creative, marketing, retail, and media companies. ![Nano Banana](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/HJbUiHvWwAA4hZn.jpeg) Google Cloud’s announcement marks a shift toward integrating AI-powered generative media into established creative and operational workflows across industries. The company worked closely with partners like Adobe and WPP, who have already embedded these models into their platforms for content production and marketing automation. Early enterprise users such as Shopify and Urban Outfitters have highlighted the models’ capacity to accelerate product development and content creation. With the introduction of video-to-image generation, Google aims to provide enterprise customers a broader set of multimodal tools to support their creative and operational goals. [Source](https://cloud.google.com/blog/products/ai-machine-learning/nano-banana-2-and-nano-banana-pro-are-generally-available?ref=testingcatalog.com) ### Anthropic launches Claude Opus 4.8 and new effort selector URL: https://www.testingcatalog.com/anthropic-launches-claude-opus-4-8-and-new-effort-selector/ Last updated: 2026-05-29T10:09:51.000Z Anthropic has released Claude Opus 4.8, the latest version of its flagship AI model, making it available to users and developers globally at the same price as its predecessor. This update improves coding, reasoning, agentic skills, and practical knowledge work, with benchmarks showing that Opus 4.8 outperforms previous Opus models and even matches or surpasses leading competitors on key tasks. Fast mode now delivers responses at 2.5 times the previous speed at a third of the former cost. > Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors. > > Available today at the same price. [pic.twitter.com/EufxL7T1kb](https://t.co/EufxL7T1kb?ref=testingcatalog.com) > > — Claude (@claudeai) [May 28, 2026](https://x.com/claudeai/status/2060042702150930686?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) New features include effort control for claude.ai, allowing users to decide how deeply the model processes tasks, and dynamic workflows in Claude Code, enabling parallel handling of very large-scale problems. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/New-chat-Claude-05-28-2026_06_49_PM--1--1.jpg) Opus 4.8 is accessible on all existing platforms, including the [Claude](https://www.testingcatalog.com/tag/claude/) website and API, for both individual and enterprise users. The new dynamic workflows feature is in research preview for Enterprise, Team, and Max plans, while effort control is available to all users. Early users and industry partners report that Opus 4.8’s judgment and reliability are improved, especially in agentic and legal applications, and that it performs more rigorous self-checks, reducing unflagged errors. The model’s alignment and safety have also undergone extensive assessment, confirming lower rates of misaligned behavior compared to earlier versions. Anthropic continues to position itself at the forefront of responsible AI development, with plans to bring even more advanced models to market in the near future. [Source](https://www.anthropic.com/news/claude-opus-4-8?ref=testingcatalog.com) ### Sesame debuts iOS app in preview with 4 personal voice agents URL: https://www.testingcatalog.com/sesame-debutes-ios-app-in-preview-with-personal-voice-agents/ Last updated: 2026-05-27T23:30:34.000Z Sesame has launched the [iOS preview](https://apps.apple.com/us/app/sesame-personal-agents/id6756329076?ref=testingcatalog.com) of its lifelike conversation partners, now accessible in the App Store across 39 countries. This rollout targets iPhone users interested in AI-driven conversational agents, with free access during the preview phase and potential waitlisting to ensure quality. The preview introduces four digital agents, Maya, Miles, Simone, and Charlie, each designed with unique personalities and voice traits to provide tailored conversational experiences. > A collection of personal agents, crafted for everyday conversation. > > Preview available now on iOS.[https://t.co/Hq5ZXAqPvX](https://t.co/Hq5ZXAqPvX?ref=testingcatalog.com) [pic.twitter.com/isDoN1bTz7](https://t.co/isDoN1bTz7?ref=testingcatalog.com) > > — Sesame (@sesame) [May 27, 2026](https://x.com/sesame/status/2059698748394262689?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) Key features include: 1. Real-time search cards with images 2. Note-taking for saving key points 3. A text mode for discreet use 4. Comprehensive memory for individualized agent experiences 5. Incognito mode for ephemeral conversations The agents are optimized for both speed and thoughtful engagement, using parallel search and retrieval systems to provide timely and accurate information. > BREAKING 🚨: Sesame just released HER > > \> Sesam iOS app is now available in Preview, offering a collection of 4 personal voice agents. > \> Sesame Agents are powered by a SOTA real-time voice mode. > \> Agents can search the web, manage reminders, and have memory. > \> App rollout is… [https://t.co/VeyPHYm3pt](https://t.co/VeyPHYm3pt?ref=testingcatalog.com) [pic.twitter.com/cFF2HaHl2V](https://t.co/cFF2HaHl2V?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [May 27, 2026](https://x.com/testingcatalog/status/2059743513047093533?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) [Sesame](https://www.testingcatalog.com/sesame-drops-first-demo-of-its-conversational-ai-voice-assistant/), the company behind this release, has focused its development on making voice-first computing accessible and intuitive. Building on feedback from an initial research preview, the team has invested in low-latency responses and improved conversational flow, distinguishing their agents from previous iterations and competitors. Early reactions from beta testers highlight the agents’ responsiveness and the natural feel of dialogues as standout qualities. ![Sesame](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/HJWd2WcbEAAS8QW--1-.jpeg) Sesame’s broader roadmap includes future Android support and the development of intelligent eyewear, signaling ongoing investment in conversational AI platforms. [Source](https://www.sesame.com/blog/voice-your-curiosity?ref=testingcatalog.com) ### Google expands Gemini for Business with shareable Projects URL: https://www.testingcatalog.com/google-expands-gemini-for-business-with-shareable-projects/ Last updated: 2026-05-27T15:44:58.000Z Google is continuing to push Gemini for Business deeper into team workflows, with several updates moving from internal development toward broader rollout. The most notable change centers on Projects, a feature that already exists in Gemini Enterprise but is being adapted for the Business tier, with a structure distinct from that of consumer Gemini. These are true container projects where individual chats live inside dedicated folders, alongside uploaded files that can be managed within the same project, a setup that turns each project into a multi-surface workspace rather than a single chat thread. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Projects-Gemini-Enterprise-05-27-2026_05_09_PM.jpg) Customization extends beyond organization. Users can assign a color to each project, define system instructions that apply across every chat inside it, and invite collaborators to work in the same workspace. That last point is where things get interesting: the collaboration model lets multiple people access and respond inside the same chat, mirroring the [group chat](https://www.testingcatalog.com/microsoft-tests-group-conversations-on-copilot/) pattern Microsoft introduced in Copilot but framed around business team contexts. This particular implementation looks unlikely to make its way to the consumer Gemini app, though further expansion across Business seats appears probable. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Agent-gallery-Gemini-Enterprise-05-27-2026_05_14_PM.jpg) In parallel, Google is bringing workflow agents to Gemini for Business, building on the agent platform already available in the Enterprise edition. A reworked builder lets users configure automated, scheduled tasks that call connectors across the Google suite and beyond, covering Gmail, Drive, Calendar, and various third-party tools. ![Gemini](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Gemini-Enterprise-05-27-2026_05_14_PM.jpg) Availability follows the usual staged pattern, with some accounts already seeing these capabilities while others wait their turn. The overall trajectory positions Gemini for Business as a closer counterpart to Enterprise, narrowing the feature gap between the two tiers while keeping shared workspaces and agent orchestration as the central pitches for paying teams. Pairing project-level memory with scheduled agents also lines [Google](https://www.testingcatalog.com/tag/gemini/) up against the always-on assistants taking shape across Microsoft, Anthropic, and OpenAI, where the competitive question is increasingly which platform can run reliable, multi-step work on behalf of a whole team rather than a single user. ### Anthropic plans expanding Claude Voice Mode to more languages URL: https://www.testingcatalog.com/anthropic-plans-expanding-claude-voice-mode-to-more-languages/ Last updated: 2026-05-27T14:40:24.000Z Anthropic appears to be preparing one of the larger updates to Claude’s mobile voice mode since its beta debut last May, with several changes surfacing in the app ahead of any official word. The refreshed UI introduces a new glow animation around the voice orb and a push-to-talk option, a meaningful tweak given that Claude’s current voice flow already follows a turn-based pattern rather than the full-duplex streaming used by ChatGPT’s Advanced Voice or Gemini Live. > ANTHROPIC 🔥: Voice mode on Claude mobile apps is about to get an upgrade with 18 new supported languages! > > \> Claude will be able to change language on the fly > \> All languages have 1-2 new voices > \> Voice Mode UI will get a new look > \> A new push-to-talk functionality will be… [pic.twitter.com/rUes3rXYdG](https://t.co/rUes3rXYdG?ref=testingcatalog.com) > > — 🚨 AI News | TestingCatalog (@testingcatalog) [May 27, 2026](https://twitter.com/testingcatalog/status/2059645806685077618?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The more consequential addition is a new “Language” setting, tagged as beta. English remains the only option live today, but the menu already lists German, Portuguese, Chinese, Japanese, Russian, Ukrainian, and other languages. Most new languages come with two voices each, while some have just one, in contrast to English, which ships with five personas (Mellow, Airy, Buttery, Glassy, and Rounded). What makes this more than a routine localization push is how the switching works. Beyond the manual selector, users can simply ask Claude to change the language mid-conversation, and it will switch on the fly, a behavior not possible in the existing build. The output still carries the cadence of text-to-speech rather than that of a native speech-to-speech model, hinting at a new orchestration layer managing multiple voices and language profiles behind the scenes, rather than a wholesale move to an in-house audio stack. That detail matters because Anthropic has so far relied on external providers for [Claude’s](https://www.testingcatalog.com/tag/claude/) spoken side, with ElevenLabs listed as a text-to-speech subcontractor and a broader relationship with Amazon powering Alexa+. A multilingual rollout layered on top of that stack would let the company close one of the more visible gaps with rivals, both ChatGPT and Gemini have offered multilingual voice for some time. No timeline has surfaced, and the language list could still shift before the public flip. ### Alook launches open-source platform for AI team orchestration URL: https://www.testingcatalog.com/alook-launches-open-source-platform-for-ai-team-orchestration/ Last updated: 2026-05-27T12:40:10.000Z [Alook](https://alook.ai/?ref=testingcatalog.com) has launched as an open-source platform that lets a single person build and direct a structured team of AI agents, coordinating them the way a founder would run a small company. The platform is now live, with the GitHub repository going public on May 25. The core concept is structural rather than technical. A user defines an org chart inside Alook, assigning each agent a role and a reporting line: dev, ops, research, writing, or whatever the project demands. From that point, work flows top-down without manual routing. A task assigned to the agent at the top is distributed automatically, with agents communicating via real email and passing deliverables down the chain, as any distributed team would. The inbox becomes the audit trail. Every instruction, reply, and handoff is recorded in email or local files, providing a complete accountability layer by default. 0:00 /1:10 1× Memory is shared across all agents in the system. No agent needs to be re-briefed on a previous decision, because every completed task feeds back into a common memory layer. The platform logs what worked and what did not after each task, using that history to build standard operating procedures that apply automatically to the next round. The intended result is compound improvement over time: the team gets faster and more accurate with every task completed, without the user having to rebuild context or re-explain conventions between sessions. The runtime runs as a persistent local daemon, meaning agents keep operating after a laptop is closed or a session ends. Users reach agents by chat or email, the same interface as any AI tool already in their workflow. The platform is agent-agnostic and works with Claude Code, Codex, and OpenCode out of the box, with more agents on the roadmap. Everything runs on the user's own machine, with full access to local tools and codebase, and no vendor lock-in in the execution path. ****SPONSORED** Test it out for yourself with Alook! [Take me there! ](https://github.com/alookai/alook?ref=testingcatalog.com) Alook positions itself in the growing category of multi-agent orchestration tools, but with a narrower, more opinionated focus than general-purpose frameworks: a single person operating a structured agent workforce locally, with email as the coordination substrate rather than API calls or visual workflow builders. The project ships as fully open-source, and the GitHub repository going live on May 25 represents the team's first broad public push after the initial launch. ### Helio launches AI-powered team workspace in Beta URL: https://www.testingcatalog.com/helio-launches-ai-powered-team-workspace-in-beta/ Last updated: 2026-05-26T23:06:43.000Z [Helio](https://bit.ly/4a4EoMC?ref=testingcatalog.com) has moved its AI Native Workforce platform to public beta with a different premise. A user describes a goal in plain language, and a built-in HR teammate translates it into a working AI team structure in under 60 seconds: the right roles, the right scope, colleagues who are live in the workspace before the conversation is over. The product is available on macOS, Windows, and via the website. ![Helio](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Screenshot-2026-05-24-at-23.54.23.webp) What happens after setup is the real argument. AI colleagues in Helio sit inside the same channels, task boards, and email threads that the rest of the team uses. They do not wait to be prompted. When a task arrives, an AI PM can break it down, assign subtasks to an AI engineer, loop in a designer, and push the whole chain forward without a human routing anything between them. AI members can also challenge each other's approach inside the same thread, surface blockers, and flag when their own reasoning is uncertain. Concrete applications the platform targets include building side projects, generating daily briefings from live data sources, reviewing contracts, monitoring competitor activity, and handling the repetitive operational work that fills up human calendars without producing much. ![Helio](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/Screenshot-2026-05-24-at-23.56.58.webp) Control stays with the team throughout. High-stakes actions, including external emails and production deploys, always route to a human approval card before anything executes. Each AI colleague runs a nightly Dream cycle, reviewing that day's conversations, identifying what worked and what did not, and updating its own working guidelines in a reversible changelog. Activity is fully traceable from the same interfaces humans use: every message, task update, and decision has a clear author and timestamp. ****SPONSORED** Test it out for yourself at Helio! [Take me there! ](https://bit.ly/4a4EoMC?ref=testingcatalog.com) Helio is the company behind the platform, based in San Francisco and founded by Wells Wang. The design argument is that AI should occupy the same organisational layer as any human colleague, with the same tasks, the same channels, and the same approval surfaces. The onboarding process for an AI teammate on Helio is designed to be substantially lighter than comparable setups on platforms like OpenClaw or Hermes Agent, requiring no terminal, no Docker, and no manual configuration. The platform integrates with Linear, GitHub, Vercel, Gmail, and Zoom, and works with Slack, Lark, Teams, and Discord as adapters for teams that already have a workspace they prefer. Helio is in public beta and accessible on macOS, Windows, and the web. Teams can join the Discord community for early access and updates. ### Anthropic to introduce personal AI Fluency scorecard in Claude URL: https://www.testingcatalog.com/anthropic-to-introduce-personal-ai-fluency-scorecard-in-claude/ Last updated: 2026-05-26T14:13:06.000Z Anthropic appears to be turning its February [research project](https://anthropic.skilljar.com/ai-fluency-framework-foundations?ref=testingcatalog.com) into a consumer-facing product. References to a new AI Fluency surface have been spotted inside Claude’s settings, where users will be able to open a dedicated screen and ask Claude to generate a personal AI fluency scorecard. The system is designed to scan a user’s activity across Chat, Cowork, and Claude Code sessions, score each session against a defined set of behavioral indicators, and produce a structured report once analysis completes, viewable and managed directly from the settings panel. > New research: The AI Fluency Index. > > We tracked 11 behaviors across thousands of [https://t.co/RxKnLNNcNR](https://t.co/RxKnLNNcNR?ref=testingcatalog.com) conversations—for example, how often people iterate and refine their work with Claude—to measure how well people collaborate with AI. > > Read more: [https://t.co/g65nGQFmjG](https://t.co/g65nGQFmjG?ref=testingcatalog.com) > > — Anthropic (@AnthropicAI) [February 23, 2026](https://twitter.com/AnthropicAI/status/2025950279099961854?ref%5Fsrc=twsrc%5Etfw&ref=testingcatalog.com) The scorecard evaluates eleven observable behaviors grouped around competencies that map closely to the [4D AI Fluency Framework](https://www.anthropic.com/research/AI-fluency-index?ref=testingcatalog.com) Anthropic built with academics Rick Dakan and Joseph Feller. The themes covered include setting the goal and approach, framing the conversation, and applying quality control, broadly the delegation, description, and discernment pillars of that framework. Early signals suggest the result is presented as a fraction, for example, 7.5 out of 11, alongside guidance on which areas a user might strengthen, giving newcomers a concrete sense of where their habits with Claude are paying off and where they aren’t. ![Claude](https://storage.ghost.io/c/2a/1b/2a1b1782-8506-4d7d-bf53-ad3fb52e2a0f/content/images/2026/05/AI-Fluency-Scorecard-----Alexey-s-Assessment-Claude-05-26-2026_01_34_AM.jpg) A sample visualisation based on the AI Fluency system prompt This is the logical next step after the AI Fluency Index Anthropic published in February 2026, which analyzed around 9,830 anonymized Claude conversations to baseline how people collaborate with AI today. That study found iteration and refinement to be the strongest predictor of good AI use, while polished outputs like artifacts and code tended to lower critical checking. Bringing the same scoring system into the product turns a research finding into a personal feedback loop, one that nudges users toward the behaviors Anthropic believes lead to safer outcomes. #### AI Fluency system prompt "Please generate a structured AI Fluency scorecard that evaluates how effectively I interact with AI across 11 behavioral indicators, based on the user messages provided below.\\n\\nThese messages are drawn from 45 conversations across 42 chat, 2 CoWork, and 1 Claude Code sessions. Each message is tagged with its surface — \[chat\], \[cowork\], or \[cc\].\\n\\nAnalyze the user messages to determine each indicator's status:\\n- Use \\"demonstrated\\" (\[+\]) for indicators where the user clearly and consistently demonstrates the skill.\\n- Use \\"partial\\" (\[\~\]) for indicators where the user sometimes demonstrates the skill or does so imperfectly.\\n- Use \\"not-observed\\" (\[-\]) for indicators where there is no evidence of the skill in the provided messages.\\n\\nFor every indicator marked \[+\] or \[\~\], include 1-2 evidence quotes taken VERBATIM from the provided messages. Keep quotes under 150 characters each. Do NOT fabricate or invent quotes — every quote must appear exactly as written in the provided messages. If a quote must be shortened to fit the limit, truncate naturally at a word boundary.\\n\\nFor every indicator (regardless of status), output a Surfaces line listing which surfaces (\[chat\], \[cowork\], \[cc\]) the supporting evidence came from. If status is \[-\], output \\"Surfaces: none\\". Fluency looks different across surfaces: coding surfaces (\[cowork\], \[cc\]) favor concise delegation; \[chat\] favors rich description. Weight Description indicators primarily against \[chat\] messages.\\n\\nBase your assessment solely on the provided messages. Do not assume skills that are not evidenced. \>>> User chat transcripts are injected here \## The 11 Indicators\\n\\nA single terse message can genuinely demonstrate multiple indicators at once. \\"ELI5\\" specifies both an audience (#2: a beginner) and a format (#3: simplified explanation). \\"less corporate\\" is both tone (#4) and implicit audience (#2). When a message packs multiple signals, credit each indicator it demonstrates — do not force it into only the single most-obvious row. The bar for each is still \\"clearly demonstrated\\", not \\"plausibly related\\".\\n\\n### Delegation\\n- 0: Clarifies goals — Does the user state what they want to accomplish before requesting help?\\n- 1: Consults on approach — Does the user ASK which approach to take before requesting execution? Interrogative: \\"what's the best way to approach this?\\", \\"how should I structure this?\\". The user is seeking a recommendation, not yet committed to a direction. Distinguish from #7: #1 asks which approach, #7 directs how Claude behaves.\\n\\n### Description\\n- 2: Defines audience — Does the user specify who the output is for?\\n- 3: Specifies format — Does the user indicate the desired output format (table, list, email, etc.)?\\n- 4: Communicates tone — Does the user indicate the voice, tone, or style they want?\\n- 5: Builds iteratively — Does the user refine outputs through follow-up rather than accepting the first result?\\n- 6: Provides examples — Does the user share examples or references to demonstrate quality expectations?\\n- 7: Sets interaction — Does the user TELL Claude how to behave, what role to adopt, or what interaction style to use? Imperative: \\"no preamble\\", \\"devil's advocate this\\", \\"steelman the other side first\\", \\"be direct\\", \\"ask me questions before writing\\". The user already knows what they want from Claude's behavior and is directing it — including when the direction is phrased as a terse request (\\"devil's advocate this\\" is role-setting, not approach-asking).\\n\\n### Discernment\\n- 8: Checks facts — Does the user question or verify factual claims in AI output?\\n- 9: Notices reasoning — Does the user push back when the AI's logic seems off? Must name a specific flaw, gap, or contradiction: \\"that doesn't follow\\", \\"you're assuming X\\", \\"that feels circular\\", \\"you skipped a step\\". Acknowledging or praising the reasoning (\\"good reasoning\\", \\"makes sense\\", \\"I follow your logic\\") does NOT count — that's acceptance, not scrutiny.\\n- 10: Recognizes context — Does the user proactively share context the AI could not know?\\n\\n## Product Feature Usage (deterministic counts from the last 30 days)\\n\\nprojects: 30 conversations (frequent)\\nartifacts: 3 conversations (sometimes)\\nweb-search: 27 conversations (frequent)\\nresearch: 3 conversations (sometimes)\\nconnectors: 4 conversations (sometimes)\\nskills: 1 conversation (sometimes)\\nmemory: 0 conversations (never used)\\nsports: 0 conversations (never used)\\nweather: 0 conversations (never used)\\nmaps: 0 conversations (never used)\\nrecipes: 0 conversations (never used)\\nsubagents: 0 conversations (never used)\\nmcp-tools: 1 conversation (sometimes)\\ncomputer-use: 0 conversations (never used)\\n\\n## Required Output Format\\n\\nOutput EXACTLY the text below — three marker-delimited sections with nothing before, after, or between them. Do NOT wrap in a code block. Do NOT add any introductory or closing text.\\n\\n--- AI Fluency Summary ---\\n\[A tight 80-110 word summary addressed directly to the user, covering BOTH collaboration behaviors and product-feature usage as one coherent paragraph. Use short, scannable sentences — no dense prose. Lead with the strongest demonstrated behavior, weave in one evidence quote, note which Claude features they rely on most, then close with one behavior and one feature to try next, framed as opportunities. Encouraging and specific, not generic.\]\\n--- End Summary ---\\n\\n--- AI Fluency Scorecard ---\\nName: User\\nRole: General\\nConversations: 45\\n\\n\[All 11 indicators in order 0 through 10\. Rules:\]\\n\[Indicator line format: \[\]