Connect with us

Technology

FuriosaAI Ends 2024 on a High Note: Llama 3.1 Performance, SDK Release, Leadership Expansion

Published

on

SANTA CLARA, Calif., Dec. 19, 2024 /PRNewswire/ — FuriosaAI, an emerging leader in AI semiconductor solutions, is closing out the year with rapid technical and customer progress with its second-generation chip, RNGD (pronounced ‘Renegade’). The recently announced AI solution has achieved compelling performance metrics in real-world enterprise deployments meeting the demand for inference with advanced large language and multimodal models.

The new performance benchmarks showcase RNGD’s ability to meet industry-leading throughput demands for Llama 3.1 models, including the 8B and 70B variants, with additional optimizations already in progress. The company also announced key software features that bring advanced optimization for customers currently sampling RNGD hardware in their production environments. These achievements represent the first phase of Furiosa’s vision for AI infrastructure that overcomes the inherent limitations of GPUs.

RNGD delivers winning throughput metrics with Llama 3.1 8B and 70B:

Building on the AI-native Tensor Contraction Processor (TCP) architecture of RNGD, Furiosa is redefining real-world AI deployments, delivering unmatched performance, programmability, and power efficiency. Furiosa’s RNGD recently achieved a throughput of 3,200–3,300 Tokens per Second (TPS) when running the LLaMA 3.1-8B model. In single-user scenarios, RNGD consistently delivers 40–60 TPS performance.

Additionally, RNGD demonstrates exceptional power efficiency, consuming 181W per card, with further optimization efforts underway. Rather than excessively boosting per-user performance, the company aims to maintain performance levels exceeding typical text-reading speeds (10–20 TPS or higher) while optimizing for multi-user environments and achieving a balanced performance approach.

Furiosa is advancing the performance and efficiency of the LLaMA 3.1-70B model. With just two RNGD cards, LLaMA 3.1-70B can be executed effectively. Currently, a single server supports up to 100 concurrent user queries, with ongoing optimizations aiming to achieve 8,000 TPS per server when equipped with 8 RNGD cards.

With the release of SDK v2024.3.0, Furiosa will expand the range of preloaded models. The SDK will also include support for tensor parallelism, enabling seamless processing across multiple elements without requiring model modifications, and a torch.compile, providing the foundation for executing customized models. Integration with HuggingFace Optimum will further empower customers to leverage a broader variety of models.

Advanced optimization tools delivered to early RNGD customers:

Building on these milestones, domestic and global enterprise customers are conducting tests with Furiosa to find a more efficient solution for scaling the inference of their self-developed models, compared to their existing setup. Their objective is to manage TCO effectively as they prepare for large-scale AI adoption. Furiosa plans to provide a high-quality AI development environment through a powerful and user-friendly SDK optimized for RNGD. The SDK v2024.1.0, currently available through the Early Access Program (EAP), is designed to handle high-performance processing of multiple LLM serving requests. It incorporates optimization techniques such as PagedAttention, Block KV Cache, and Continuous Batching, while also supporting various token sampling methods, including Greedy, Beam Search, and Top-k/p. These features allow developers to seamlessly create AI services customized to meet a wide range of requirements. The SDK and online sample will be available after the release of v2024.3.0.

Furiosa remains committed to delivering the most sustainable AI deployment solutions through rigorous optimization at an unprecedented pace.

“With RNGD now in customers’ hands, we are accelerating the next generation of frontier LLMs to unlock emerging Agentic AI applications—bringing advanced reasoning capabilities to enterprise verticals, all at dramatically lower costs,” said June Paik, Co-Founder and CEO of FuriosaAI.

Furiosa Expands Global Footprint with Strategic Leadership Appointment

Furiosa is scaling production and expanding its leadership team with the appointment of Alex Liu as Senior Vice President of Product and Business. A Technology Emmy Award winner and co-founder of NETINT Technologies, Alex brings over 20 years of expertise in startup management, technology innovation, and strategic leadership. At NETINT, he spearheaded groundbreaking achievements, including the development of the world’s first VPU SoC, setting new industry benchmarks and securing the prestigious 2024 Technology Emmy Award. At Furiosa, Alex will lead global product management, go-to-market strategies, and partnerships to drive innovation and align the company’s AI-native technologies with a vision to empower the development of planet-scale AI infrastructure.

RNGD is currently sampling with customers, and mass production will ramp up in partnership with TSMC for 2025 availability. To learn more about Furiosa, please visit https://furiosa.ai/.

About FuriosaAI

FuriosaAI is a semiconductor company dedicated to creating sustainable AI computing solutions that make powerful AI accessible to all. With its innovative Tensor Contraction Processor architecture, FuriosaAI is revolutionizing the AI hardware landscape, offering unparalleled efficiency and programmability for the most demanding AI workloads. For more information, please visit https://furiosa.ai/.

View original content to download multimedia:https://www.prnewswire.com/news-releases/furiosaai-ends-2024-on-a-high-note-llama-3-1-performance-sdk-release-leadership-expansion-302336756.html

SOURCE FuriosaAI

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Technology

Yiren Digital Accelerates Operating Efficiency Through AI Agent Deployment

Published

on

By

Broader AI adoption improves productivity across asset recovery and enterprise operations

BEIJING, July 23, 2026 /PRNewswire/ — Yiren Digital Ltd. (NYSE: YRD) (“Yiren Digital” or the “Company”), a leading company specializing in financial technology and artificial intelligence innovation across multiple industries in China and global markets, today announced measurable operating efficiency improvements as it continues to deploy AI agents across core enterprise workflows. Broader AI adoption is reducing manual intervention, increasing workforce productivity and creating greater operating leverage by automating high-volume processes across multiple business functions.

These deployments are a key component of Yiren Digital’s “All-in-AI” strategy and its broader transition from AI-assisted productivity toward agent-driven execution. By embedding AI agents into core workflows, the Company is creating reusable operating capabilities that can be deployed across its businesses, supporting greater efficiency and reducing the cost of extending automation into new functions.

“Our objective is not simply to automate individual tasks, but to fundamentally improve how work is performed across the enterprise,” said Mr. Ning Tang, Chairman and Chief Executive Officer of Yiren Digital. “As AI agents take on more of our high-volume, demanding workflows, the productivity gains are becoming a structural part of how we run the business, not a one-time efficiency project. We will continue to deepen AI integration across our existing businesses while extending reusable capabilities into additional verticals.”

The AI deployments are supported by the Company’s proprietary enterprise AI architecture, including MagiCube 2.0, its upgraded multi-agent platform. The platform provides common infrastructure for agents deployed across marketing, customer service, capital operations, risk management, compliance and research and development, with more than 10 reusable foundational capabilities, supporting enterprise-wide execution.

Measurable Operating Impact

Lower manual intervention: The human handling rate in asset-recovery operations decreased from 45.0% to 24.9%, representing a 20.1-percentage-point decline, an approximately 44.6% relative reduction in manual intervention.

Higher staff productivity: The number of service tickets handled per asset-recovery staff member within the applicable Month 1 workflow increased from 358 to 525, an improvement of approximately 47%.

Expanded agent adoption: AI agents accounted for 81% of service tickets within eligible Day 1 asset-recovery workflows in 2025, up from 50% in 2024. The Company also deployed AI agents selectively in later-stage workflows, accounting for 20% of eligible service tickets at Day 4, 14% at Day 16 and 20% at Month 2. Each percentage is calculated separately for the relevant stage and should not be interpreted as a sequential adoption trend.

Enterprise-wide reuse: MagiCube 2.0 supports agent deployment across six enterprise functions, allowing the Company to apply common AI capabilities to a broader range of regulated and high-volume workflows.

Enterprise-scale AI execution: The Fengchao AI voice agent processes approximately 1,500 hours of real-time speech-to-text activity each day. The LingShu intelligent marketing platform executes more than 1,700 tasks daily and generates individualized communication content in an average of 0.6 seconds.

Building Enterprise Operating Leverage Through AI

As AI deployment expands across the enterprise, Yiren Digital is increasingly shifting repetitive, high-volume tasks from human-assisted processes toward agent-driven execution. By combining AI agents with centralized orchestration and governance, the Company is improving operating consistency, strengthening workforce productivity and creating reusable capabilities that increase operating leverage as AI is deployed across additional business functions.

Yiren Digital plans to continue expanding agent-driven workflows across its credit and insurance operations, as part of its ongoing All-in-AI strategy, while strengthening the shared architecture and governance that support enterprise-wide AI deployment. These capabilities are designed to scale across multiple use cases and provide a foundation for the Company’s broader expansion into AI application-layer opportunities, including AI entertainment and AI-assisted language learning.

About Yiren Digital

Yiren Digital Ltd. is a leading company specializing in financial technology and artificial intelligence innovation across multiple industries in China and global markets. The Company leverages advanced artificial intelligence and emerging technologies to enhance customer experience, optimize capital efficiency, and expand financial inclusion. Following the regulatory filing of its in-house developed Large Language Model Zhiyu, and the significant enhancement of its MagiCube Agent platform, Yiren Digital is establishing a new growth engine to accelerate its evolution into an AI-native, multi-industry operating platform extending beyond traditional financial services. For more information, please visit https://ir.yiren.com.

Safe Harbor Statement

This press release contains forward-looking statements. These statements are made under the “safe harbor” provisions of the U.S. Private Securities Litigation Reform Act of 1995. These forward-looking statements can be identified by terminology such as “aim,” “anticipate,” “believe,” “estimate,” “expect,” “hope,” “going forward,” “intend,” “ought to,” “plan,” “project,” “potential,” “seek,” “may,” “might,” “can,” “could,” “will,” “would,” “shall,” “should,” “is likely to” and the negative form of these words and other similar expressions. This press release contains forward-looking statements within the meaning of Section 21E of the Securities Exchange Act of 1934, as amended, and as defined in the U.S. Private Securities Litigation Reform Act of 1995. These statements can be identified by terminology such as “will,” “expects,” “anticipates,” “future,” “intends,” “plans,” “believes,” “estimates,” “target,” “confident,” and similar expressions. Forward-looking statements are based on management’s current expectations, assumptions, and assessments of current market and operating conditions. These statements involve inherent risks, uncertainties, and other factors, many of which are outside the control of the Company, and which could cause actual results to differ materially from those expressed or implied in such statements. Actual results may differ materially from those expressed or implied in forward-looking statements due to a variety of factors and other risks described in the Company’s filings with the U.S. Securities and Exchange Commission. All forward-looking statements speak only as of the date of this press release. The Company undertakes no, and expressly disclaims any, obligation to update or revise any forward-looking statements, whether as a result of new information, future events, or otherwise, except as required under applicable law.

View original content:https://www.prnewswire.com/news-releases/yiren-digital-accelerates-operating-efficiency-through-ai-agent-deployment-302833201.html

SOURCE Yiren Digital Ltd.

Continue Reading

Technology

Infinium Edge Launches EdgeSites™, a New Infrastructure Model for Deploying AI Compute at Existing Commercial and Industrial Facilities

Published

on

By

EdgeSites delivers operational AI infrastructure in existing powered buildings — factory-built data center modules, waterless cooling, and ready in months without new construction or grid interconnection required.

SACRAMENTO, Calif., July 23, 2026 /PRNewswire/ — Infinium Edge™ today announced Infinium EdgeSites™, a development program that utilizes existing commercial and industrial facilities to deploy operational AI compute infrastructure. Built around Infinium Edge’s proprietary Edge Thermal Vectoring™ immersion cooling platform, EdgeSites enables high-density GPU deployments in existing buildings that were never designed as data centers — without new construction, without cooling water infrastructure, and without the multi-year grid interconnection timelines that constrain conventional large-scale data center development.

More than 20 million commercial and industrial electricity customers in the US are served by electrical infrastructure sized to peak demand – which industry research shows are utilized at only 40-60% on average. That unused headroom, capacity already contracted, energized, and sitting behind the meter, can support high-density AI compute without adding new load to the grid or waiting on a new interconnection.

At the center of the program is the Vector ONE™ — Edge’s factory-built, self-contained immersion cooling system designed to house 1 MW of AI compute capacity. Vector ONE units are engineered for deployment in standard commercial and industrial buildings, either indoors or outdoors, arriving pre-integrated, fully commissioned and require no municipal water connection. Installations are modular and scalable: additional units can be commissioned as site power and demand allow, without rebuilding the underlying infrastructure and occupy up to 70% less floor space than air-cooled equivalents.

Built for the Shift to Inference

As inference moves to displace training as the dominant AI workload, the growth opportunity is shifting towards small, distributed data centers that can be deployed quickly and sited where demand originates. Conventional data center developments are under compounding pressure from long utility interconnection queues, sometimes lasting years, pressure around water use, and general community and regulatory opposition enacting restrictions. Community opposition and regulatory friction delayed or blocked an estimated $156 billion in planned U.S. data center capacity in 2025 alone.

EdgeSites is purpose-built for the structural shift to inference and addresses key issues stalling conventional data center developments today. Each Vector ONE unit delivers 1 MW of inference-ready capacity inside an existing building, in a market that already has established electrical infrastructure, in a timeline measured in months rather than years. Multiple units can be used in tandem to deploy up to 10 MW of capacity at a single site.  The program converts the distributed inventory of underutilized industrial or commercial electrical capacity in the United States into a nationally scaled inference network. Vector ONE’s dry-cooler loop consumes no municipal water, making EdgeSites viable in markets where evaporative cooling has been restricted or banned.

“The data center industry has been answering an infrastructure shortage with a construction playbook — build new facilities, secure new grid connections, wait years for capacity to come online,” said Robert Schuetzle, CEO of Infinium. “That model cannot keep pace with AI deployment timelines. Infinium EdgeSites operate around different premises: the power already exists, the buildings already exist, and the technology now exists to put them to work. We are making operational what the industry has been treating as stranded.”

Deploying EdgeSites

As demand for AI compute continues to outpace available infrastructure and focuses on distributed inference needs, Infinium Edge is expanding the EdgeSites network with qualified host locations and compute partners.

Commercial and industrial property owners of industrial sites, distribution centers, warehouses, or large commercial properties with available electrical capacity benefit from receiving lease income from infrastructure they already own or control. Infinium Edge manages all aspects of site development and operations for installing and deploying the Vector ONE system. No capital investment or operational responsibility is required from the host.

AI companies, enterprises, and compute operators requiring infrastructure on compressed deployment timelines can access high-density, edge-proximate GPU capacity through a straightforward capacity agreement, priced by the kilowatt-month, with backup power included in the capacity fee. There is no construction to manage, no permitting process to navigate, and no cooling infrastructure to operate or maintain.

Infinium Edge manages the full program from development and installation to operation and monitoring— simplifying development and data center management for AI companies and enterprises.

Reach out to learn more and partner in EdgeSites deployments.

Inquiries: www.infinium.ai/edgesites

About Infinium Edge™
Infinium Edge™ is the advanced AI data center infrastructure platform from Infinium, delivering high-density, sustainable compute through proprietary single-phase immersion cooling technology. Infinium Edge is the only North American producer of Fischer-Tropsch immersion fluids and offers a full-stack platform — including Edge Thermal Vectoring™ platform, Vector ONE™ modular AI Factory units, ETV100 immersion fluids, and integrated monitoring systems — engineered for the thermal and operational demands of AI and high-performance computing at scale. For more information, visit www.infinium.ai.

View original content to download multimedia:https://www.prnewswire.com/news-releases/infinium-edge-launches-edgesites-a-new-infrastructure-model-for-deploying-ai-compute-at-existing-commercial-and-industrial-facilities-302832792.html

SOURCE Infinium

Continue Reading

Technology

ChipMOS SCHEDULES SECOND QUARTER 2026 FINANCIAL RESULTS SEMIANNUAL CONFERENCE CALL

Published

on

By

HSINCHU, July 23, 2026 /PRNewswire-FirstCall/ — ChipMOS TECHNOLOGIES INC. (“ChipMOS” or the “Company”) (Taiwan Stock Exchange: 8150 and Nasdaq: IMOS), an industry leading provider of outsourced semiconductor assembly and test services (“OSAT”), today announced that it will report second quarter 2026 results and host a semiannual conference call after the close of trading on the Taiwan Stock Exchange on Tuesday, August 11, 2026.

Investors and analysts are encouraged to participate in the semiannual conference call using the dial-in phone number noted below. A webcast and replay will be available on the Company’s website.

Date: Tuesday, August 11, 2026
Time: 3:00PM Taiwan (3:00AM New York)
Dial-In: +886-2-3396 1191
Password: 1637011 #

Semiannual Conference Call Webcast and Replay: https://www.chipmos.com/chinese/ir/info2.aspx
Replay: Starts Approximately 2 hours after the live call ends

Language: Mandarin

Note: A transcript will be provided on the Company’s website in English following the semiannual conference call to help ensure transparency, and to facilitate a better understanding of the Company’s financial results and operating environment.

About ChipMOS TECHNOLOGIES INC.:
ChipMOS TECHNOLOGIES INC. (“ChipMOS” or the “Company”) (Taiwan Stock Exchange: 8150 and Nasdaq: IMOS) (www.chipmos.com) is an industry leading provider of outsourced semiconductor assembly and test services. With advanced facilities in Hsinchu Science Park, Hsinchu Industrial Park and Southern Taiwan Science Park in Taiwan, ChipMOS is known for its track record of excellence and history of innovation. The Company provides end-to-end assembly and test services to leading fabless semiconductor companies, integrated device manufacturers and independent semiconductor foundries serving virtually all end markets worldwide.

Forward-Looking Statements:
This press release may contain certain forward-looking statements. These forward-looking statements may be identified by words such as ‘believes,’ ‘expects,’ ‘anticipates,’ ‘projects,’ ‘intends,’ ‘should,’ ‘seeks,’ ‘estimates,’ ‘future’ or similar expressions or by discussion of, among other things, strategies, goals, plans or intentions. These statements may include financial projections and estimates and their underlying assumptions, statements regarding current macroeconomic conditions, including the impacts of high inflation, foreign exchange rates and risk of recession, on demand for our products, consumer confidence and financial markets generally; changes in trade regulations, policies, and agreements and the imposition of tariffs that affect our products or operations, including potential new tariffs that may be imposed and our ability to mitigate with respect to future operations, products and services, and statements regarding future performance. Actual results may differ materially in the future from those reflected in forward-looking statements contained in this document, based on a number of important factors and risks, which are more specifically identified in the Company’s most recent U.S. Securities and Exchange Commission (the “SEC”) filings. Further information regarding these risks, uncertainties and other factors are included in the Company’s most recent Annual Report on Form 20-F filed with the SEC and in its other filings with the SEC.

Contacts:

In Taiwan

Jesse Huang

ChipMOS TECHNOLOGIES INC.

+886-6-5052388 ext. 7715

IR@chipmos.com

In the U.S.

David Pasquale

Global IR Partners

+1-914-337-8801

dpasquale@globalirpartners.com

 

View original content:https://www.prnewswire.com/news-releases/chipmos-schedules-second-quarter-2026-financial-results-semiannual-conference-call-302831885.html

SOURCE ChipMOS TECHNOLOGIES INC.

Continue Reading

Trending