The Importance Of Security Benchmarking When Using AI Development Tools - TalkLPnews Skip to content

The Importance Of Security Benchmarking When Using AI Development Tools

The Importance Of Security Benchmarking When Using AI Development Tools
The Importance Of Security Benchmarking When Using AI Development Tools

GUEST OPINION:  If it was possible to have an experienced chef in the kitchen to help prepare a sophisticated dish, most people would happily accept the assistance. The same holds true for a qualified contractor to work on a home-improvement project, or an office aide to handle tedious, repetitive tasks.

For this reason, it comes as little surprise that software developers are rapidly adopting large language model (LLM) tools and additional forms of generative artificial intelligence (Gen AI) to help them produce code more efficiently, and at a rapid pace.

During the past five years, nearly two-thirds of IT executives and administrators say their organisation has incorporated these AI tools into the software development lifecycle (SDLC), according to research from KPMG and OutSystems, which makes available a low-code, AI-supported platform to build applications.

Growing numbers of available tools

Three-quarters of those surveyed say 10 to 50% of code in final products is created with Gen AI technologies, and more than nine of ten intend to boost AI investments further. They’re doing so because of the clear benefits, as four of five say they are seeing as much as a 50% reduction in development time due to the greater usage of AI and automation tools.

With constant newcomers to the market promising better productivity and output than the last, the increased ubiquity of AI in the SDLC appears inevitable. However, 56% of these executives cite data privacy and security concerns as the main barriers to adoption.

The challenges of effective tool assessment

The concerns are valid. Even if tools state they have “improved” protection in new versions, it cannot be assumed they are secure by default. BaxBench, which oversees a coding benchmark to evaluate LLMs for accuracy and security, has concluded that no current LLM can generate deployment-ready code. It indicates that 62% of solutions produced by even the best model are either incorrect or contain a vulnerability. Among those that are correct, about one-half are insecure.

Often readily exploitable and prone to insecure code output, the tools are very likely to trigger compromises. China’s DeepSeek, for example, has emerged as a popular option, with between 5 and 6 million users worldwide.

It presents a faster and smarter assistant for software development, compared to well-established LLMs. It’s also much more affordable, priced at around 1/30th of the cost of similar models.

Yet, research has shown that DeepSeek is susceptible to critical risks, such as malware generation (with a failure rate of 93%), jailbreaking (91%) and prompt injection attacks (86%). This performance indicates that DeepSeek is too unsafe for business and enterprise use.

Regardless, organisations often continue to use these tools, frequently without the knowledge or approval of executives. This is commonly referred to as “shadow AI” which is a highly concerning issue.

A standard way of measuring AI tool security

Given the continuously increasing deployments coupled with precarious protection, security leaders and development teams – and the industry as a whole – must come together to establish a standard for using AI coding tools.

Currently, there is no uniform process that teams can easily follow to assess products for safety or compare them to other tools. Everyone is proceeding in different ways, and this piecemeal approach can introduce risk in the form of vulnerabilities.

So what should standardisation look like? It starts with benchmarking, which focuses on two key areas:

Tools: Security leaders and development team personnel need to assess the data sources feeding an LLM tool, and how the tool generates code. They should also identify the mechanisms in place that either will (or will not) keep cyber criminals from exploiting the code. Then, they must compile individual scores for each of these considerations, and combine those to come up with an overall security score.

People: Tools are only one-half of the solution while developers are the other. It’s essential for them to learn and upskill in security so they can make informed decisions about AI usage that better protect code, as opposed to exposing it. Toward this goal, their organisations have to invest in education for developers so they can write safeguarded code from the start while mitigating vulnerabilities, including those generated by AI coding assistants.

Learning to code correctly with safe coding patterns via additional agile, hands-on learning pathways proves to be the most effective approach. It helps developers learn, test and apply knowledge immediately and with real-world context, and then come up with defence approaches that emerge as second-nature routines.

Again, standardisation benchmarks tracking training frequency/impact, team skill levels and vulnerability reductions will enable organisations to measure – and improve – enterprise-wide security maturity.

AI-powered tools will continue to provide significant support to software developers. However, the need for human oversight remains to ensure the effective quality and security is maintained at all times.

http://itwire.com/guest-articles/guest-opinion/the-importance-of-security-benchmarking-when-using-ai-development-tools.html